Roadmap · beyond v2026.07

The run scores 7 of the 17 stages of go-to-market work.

Eight task families make a benchmark. Go-to-market is seventeen stages. This run reaches stages 1–11: 6 scored on their own terms, 2 touched in passing, 9 absent altogether. Nothing after stage 11 is measured at all, which means the entire post-pipeline motion sits outside the score. This page is that audit, and what a second version would have to do about it.

Lifecycle coverage

Every stage of the revenue motion, and the family that scores it, or the admission that nothing does.

41%weighted coverage
partial credit at 0.5

  1. 01Company understandingScored · Company & persona researchScored as a cited dossier, with a hard penalty for inferred scope.
  2. 02ICP reasoningNot measuredNothing asks a model to infer an ICP from won/lost history and defend it.
  3. 03Market researchPartial · Account researchAccount research surfaces market context but is graded as account context, not market analysis.
  4. 04Buyer researchScored · Company & persona researchBuying-committee coverage is an explicit rubric dimension.
  5. 05Competitive intelligenceNot measuredCompetitors appear only as a scripted objection, never as an analysis task.
  6. 06Prospect generationNot measuredAccounts are handed to the model. Nothing scores finding them, or the precision of a built list.
  7. 07Message creationScored · Outbound copyFirst touch and follow-up, against a spam classifier and a specificity rubric.
  8. 08Objection handlingScored · Objection handlingSix turns against a scripted buyer, scored on holding price and finding the real blocker.
  9. 09Multi-channel sequencingScored · Sequence executionFourteen days, eight tools, scored on plan coherence and recovery.
  10. 10CRM updatesScored · CRM extractionSchema-strict extraction with deduplication; a schema violation scores zero.
  11. 11Follow-upPartial · Sequence executionFollow-up is planned inside sequence execution, never scored against a real reply.
  12. 12Meeting preparationNot measuredCall analysis grades the debrief. Nothing grades the brief a rep reads going in.
  13. 13Pipeline managementNot measuredNo stage hygiene, no next-step discipline, no multi-deal triage.
  14. 14ForecastingNot measuredNo commit/best-case calls, and no scoring against what actually closed.
  15. 15Deal coachingNot measuredNothing asks what is wrong with this deal and what the rep should do on Monday.
  16. 16Renewal & expansionNot measuredThe whole post-close motion is out of scope.
  17. 17Customer successNot measuredHealth, risk and adoption are unmeasured.

Two families that do not map to a stage

  • Call analysis: Post-call synthesis, sitting between stages 11 and 12 without being either.
  • End-to-end account plan: Spans stages 1–9 as one artifact, which is why it is the hardest family in the run.

The five gaps

Coverage is the easy problem. The hard one is that the tasks are the wrong shape.

Adding nine more families would widen the benchmark without changing what it measures: a single-shot answer to a well-formed question, scored on whether the answer was right. Production go-to-market is agentic, continuous, priced in latency as much as tokens, and fed data that contradicts itself. Those are different measurements, not more of the same one.

The five measurement gaps

Gap 01

Tool use, and recovering when it fails

Sequence execution is the only family with tools, and it grades the plan more than the calling. A GTM agent in production decides which tool to reach for, merges evidence from several that disagree, and keeps going when one returns nothing. None of that is scored today.

  • Tool selection: was the right call made, or the expensive one?
  • Evidence merge: two sources conflict. Which wins, and is the choice defended?
  • Failure recovery: an enrichment call 404s mid-run. Retry, route around, or stop?
  • Abstention: no data found. Does the model say so, or invent a plausible contact?

Gap 02

Accuracy is one of three axes

The run scores whether the answer was right and publishes what the tokens cost. It does not measure how long the answer took or how much work it took to get there, and two systems that both solve the task can be entirely different products.

  • Model A: 45s, 20 tool calls, $0.60
  • Model B: 6s, 3 tool calls, $0.04
  • Same score today. Not the same system in production.
  • Wall-clock latency, p95 as well as median, is the missing column.

Gap 03

Why this account, and why not that one

Grading whether a model found the right company rewards the answer and ignores the reasoning that produced it, which is exactly backwards from how an enterprise buyer evaluates the system. The rejected candidates carry as much signal as the chosen ones.

  • Justification: why this account, in terms a rep can repeat to their manager.
  • Counterfactual: why the near-miss was rejected.
  • Citation quality: does the source actually support the claim, or merely mention the company?
  • Confidence calibration: does stated confidence track accuracy, or is everything asserted flat?

Gap 04

The benchmark stops at the recommendation

Eight families each grade one artifact. A GTM hire is not judged on eight artifacts in isolation. They are judged on whether one named deal moved. The unit of evaluation should be an account, not a task.

  • Input: "Sell Salestools AI into Snowflake."
  • Output: research, ICP fit, buyer map, first-touch email, LinkedIn message, follow-up sequence, objection prep, CRM entry, meeting agenda, account plan.
  • Scored as one deliverable, including whether the parts contradict each other.
  • Internal consistency is the thing eight separate scores cannot see.

Gap 05

Enterprise data is not clean, and the tasks are

Every task in the run is well-formed. Real accounts have been acquired, have three websites that disagree, have duplicate contacts and stale titles, and sit under compliance rules that forbid the obvious play. Ambiguity should be injected deliberately, because how a model behaves when the data is bad is the whole question.

  • Conflicting sources: the site says 200 staff, the filing says 1,400.
  • Corporate structure: the buyer works at a subsidiary under a different name.
  • Stale data: the champion left four months ago.
  • Compliance: the obvious channel is not available in this jurisdiction.
  • The score is not whether the model copes. It is whether it asks, flags uncertainty, or fabricates.

The measurement this design cannot express

Every task in the run begins and ends inside one context window. An account does not.

  1. Day 1Agent discovers the account and builds the initial plan.
  2. Day 7The company raises a Series C.
  3. Day 10The CTO, the agent's primary persona, is replaced.
  4. Day 15A competing product launches into the same category.
  5. Day 18The buyer finally replies, to a thread from Day 2.
  6. Day 19Does the strategy change, and is the change defensible?

Scoring this means holding state for nineteen days, re-deciding on new evidence, and grading a revision rather than an answer: whether the strategy changed when it should have, and stayed put when it should not. Nothing in a single-shot, temperature-0, no-retries harness can reach it. It is the reason the eight families are a floor and not a finish line.

What lands first

Latency and tool-call counts are already emitted by the harness and simply are not published. That is a reporting change, not a research one, and it is the next column on the board. Reasoning quality and injected ambiguity need new rubrics but reuse the existing grader pipeline. End-to-end execution and continuous agents need a different harness altogether.

← Back to the 30-model leaderboard

Proposed v2 composite

GTM-2  =  quality  ×  efficiency  ×  reliability

quality     rubric score, as today
efficiency  f(tokens, tool calls, wall-clock p95)
reliability abstention + calibration + recovery

Reported as three numbers, never collapsed
into one. A system can be excellent on any
two and unusable on the third.