Nobody Independent Is Checking If AI Actually Works [Sources]

Nobody Independent Is Checking If AI Actually Works [Sources]

udipta basumatari

# Mispriced - "AI's Productivity Is Mispriced" · Reference Sheet


Every claim below maps to the section it appears in. Provenance flags:

`[vendor-reported]` maker's own figure · `[interpolated]` line drawn between sourced endpoints · `[confirm at publish]` fast-moving or contributor-sourced, re-check the primary before publishing.


---


## Hook - Salesforce pay-per-resolution


- Agentforce pay-per-resolution, ~$2 per successful resolution, no charge on human escalation or customer walk-away, announced June 25, 2026, GA July 2026.

 - Salesforce Ben: https://www.salesforceben.com/huge-agentforce-pricing-shift-salesforce-introduces-pay-per-resolution/

 - Constellation Research: https://www.constellationr.com/insights/news/salesforce-takes-run-outcome-based-help-agent-pricing

 - Portal ERP (resolution defined by rules, 1,000-unit blocks, 10-min window): https://portalerp.com/article/salesforce-introduces-service-agent-with-outcome-based-pricing-structure

 - The Leveraged Years (CIO-sourced, $2/resolution): https://www.theleveragedyears.com/ai-workflows/salesforce-agentforce-pay-per-resolution-pricing


- Pricing history: $2/conversation (2024) → $0.10/action Flex Credits (May 2025) → $125/user/mo (2025) → $2/resolution (June 2026).

 - SaaStr (all three models, timeline, $540M ARR / 330% YoY, ~8% adoption): https://www.saastr.com/salesforce-now-has-3-pricing-models-for-agentforce-and-maybe-right-now-thats-the-way-to-do-it/

 - Salesforce press release (Flex Credits, $2/conversation history): https://www.salesforce.com/news/press-releases/2025/05/15/agentforce-flexible-pricing-news/

 - Getmonetizely (pricing whiplash timeline): https://www.getmonetizely.com/blogs/the-doomed-evolution-of-salesforces-agentforce-pricing

 - Chart 2 source. Solid, multi-source.


## The cost panic - Uber, Nvidia, OpenAI


- Uber burned its entire 2026 AI coding budget in 4 months; 84% of engineers on Claude Code by March; ~70% of committed code AI-originated; president said token use didn't correlate with shipped features. `[confirm at publish]`

 - Forbes (Jemma Green, primary framing): https://www.forbes.com/sites/jemmagreen/2026/07/02/ai-costs-more-than-the-people-it-replaced/

 - Yahoo Finance (Microsoft + Uber AI budget): https://ca.finance.yahoo.com/news/microsoft-uber-hit-ai-budget-113000485.html

 - Fast Company (corroborates Uber budget, Catanzaro quote): https://www.fastcompany.com/91568851/the-career-edge-that-no-algorithm-can-take-from-you

 - Note: Forbes is a contributor column; the exec statements are quoted there and corroborated by Fast Company + Yahoo. Re-confirm the exact 4-month / 84% / 70% figures against an Uber exec's own words or a wire report before publish.


- Nvidia VP of applied deep learning (Bryan Catanzaro): compute for his team now exceeds spend on the employees using it. `[confirm at publish]`

 - Forbes (above); Fast Company (above).


- OpenAI spends ~$2 per $1 earned on inference; loses money on $200/mo subscriptions (Altman). `[confirm at publish]`

 - Forbes (above).

 - Yahoo Finance (Altman "losing money"): https://finance.yahoo.com/news/sam-altman-says-losing-money-080700756.html


## "Costs are falling" is a mirage


- Gartner: running the largest models could be ~90% cheaper by 2030. `[confirm at publish]`

 - Forbes (cites Gartner). Find the primary Gartner release before publish.

- Consumption scaling faster than price falls; Amazon usage leaderboard gamed and pulled.

 - Forbes (above).

- Chart 3 is `[interpolated]`: only endpoint (Gartner ~90% by 2030) is sourced; both trend lines are illustrative, labeled ILLUSTRATIVE on the chart face. Do not present as measured data.


## Self-graded outcomes - resolution definitions


- Salesforce CRMArena-Pro benchmark: out-of-box Agentforce agent ~35% accuracy before customization. `[vendor-reported]`

- Reported "resolution" rates, mixed definitions: GE Appliances 25%, Reddit 46%, Pandora 60%, OpenTable 70%, Salesforce Customer Zero/help portal 77% (also cited 70%), 1-800Accountant 90%, Heathrow ~95%. `[vendor-reported]`

 - My AskAI (compiles all of the above, notes the inconsistent definitions): https://myaskai.com/blog/salesforce-agentforce-complete-guide-2026

 - Chart 4 source. Flagged vendor-reported on the chart face.


## Kimi K3 (mirage section closer + self-grading opener)


- Moonshot AI released Kimi K3, the largest open-weight model shipped to date, claiming frontier-competitive performance at a fraction of the cost. Rival AI stocks fell sharply the following day; Fortune framed it as a second DeepSeek shock. `[vendor-reported claims, independent market reaction]`

 - Fortune (release, "competitively" with Fable 5, largest open-weight model): https://fortune.com/2026/07/16/moonshots-kimi-k3-pushes-chinese-ai-into-fable-level-territory/

 - Fortune (DeepSeek-shock framing, "substantially outperformed" Opus 4.8 and GPT 5.6 Sol, per Moonshot): https://fortune.com/2026/07/17/china-moonshot-kimi-k3-markets-china-ai/

 - Bloomberg (K3 outperforms all rivals except Claude Fable 5 and GPT-5.6, "it said"): https://www.bloomberg.com/news/articles/2026-07-17/china-s-powerful-new-moonshot-ai-model-closes-gap-with-us-rivals

 - Quartz via Yahoo Finance (below Fable 5 and GPT 5.6 Sol overall; beat Opus 4.8 and GPT 5.5 on coding/agent evals; rival AI stocks lower): https://finance.yahoo.com/technology/ai/articles/moonshot-ai-kimi-k3-launch-122826603.html

 - TechCrunch (pre-release, FT-sourced): https://techcrunch.com/2026/07/16/moonshots-upcoming-kimi-3-is-expected-to-close-the-gap-with-anthropics-opus-4-8/


- THE LOAD-BEARING POINT, and it is a contradiction in the coverage, not in our analysis: Quartz places K3 BELOW GPT-5.6 Sol overall, while Fortune reports Moonshot claiming K3 SUBSTANTIALLY OUTPERFORMED GPT-5.6 Sol. Same named model, opposite direction, inside 24 hours, both tracing to the vendor's own materials. The article asserts no ranking of its own. `[confirm at publish]`

 - Before publishing, re-read Moonshot's official K3 tech blog and confirm the discrepancy is genuine rather than an artifact of different variants or benchmark subsets. If it resolves cleanly, soften to "the coverage reported the claim inconsistently" rather than dropping the beat, since the self-reporting point stands either way.


- DATE DISCREPANCY `[confirm at publish]`: Fortune says the unveiling was July 16, 2026; Bloomberg says Moonshot released it Friday, which was July 17, 2026. The official tech blog appears to have posted July 17. The article uses July 16, 2026 for the release and "the following day" for the market reaction. Verify against Moonshot's own post.

- PARAMETER COUNT `[confirm at publish]`: reported as 2.7 trillion (Fortune) and 2.8 trillion (Quartz, explainX). The article avoids the number entirely and says "largest open-weight model shipped to date," which every source agrees on.

- DELIBERATELY OUT OF SCOPE: Anthropic accused Moonshot of distillation campaigns in February 2026, and there is open debate about whether K3 closed the gap by distilling frontier models. The article does not touch this. Dumpo has an Anthropic-assistance disclosure in the piece and no independent way to adjudicate the claim, so entering that fight would trade a defensible thesis for an indefensible one.

- ALSO OUT OF SCOPE, HELD FOR A FUTURE PIECE: cheap frontier-adjacent open weights are the case study for the token-cost-collapse and cheap-inference-fallacy angles. This piece spends only the two sentences needed to pre-empt the "but it just got cheap" rebuttal.


## Databricks internal benchmark (self-grading section + the close)


- Databricks, July 8, 2026, built an internal coding-agent benchmark from its engineers' own merged pull requests on a multi-million-line codebase, graded against the team's own tests. INDEPENDENT of the model vendors, though it is the company's own workload and Databricks sells the routing and gateway layer, not the models. Not a generalizable public benchmark and the article does not present it as one.

 - https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase

 - Why they rejected public benchmarks: the tasks are public, so solutions leak into training data over time, and results were not representative of a codebase spanning 10+ languages.

 - No LLM judge, explicitly because it rewards sounding right over being right.

 - Sealed git history mid-run after finding agents with shell access could walk forward through commits to recover the merged solution.

 - Per-token price is a poor predictor of per-task cost: Sonnet 5 is ~1.7x cheaper per token than Opus 4.8, yet cost $2.09/task against Opus's $1.94 while scoring 6 points lower on completion (81% vs 87%), because it read more and consumed 1.9x the tokens. The article keeps these figures but does not name the models, per Dumpo's framing preference; the names are here if he wants them added.

 - Harness choice alone swung cost per task by more than 2x at identical quality.

 - Also in the post but not used: GLM 5.2 landed in the top capability tier, statistically tied with Opus 4.8 on quality at $1.28/task against $1.94. This is strong material for the held open-weights/token-cost piece.

 - Used twice: to introduce the benchmark-contamination point in the self-grading section, and as the positive counterexample in the close. It is the piece's proof that this is not an anti-AI argument.


## Self-graded efficiency - the frontier launches (supporting pattern in the self-grading section)


- Two flagship models launched a day apart, both leading with efficiency rather than raw capability: Grok 4.5 (July 8, 2026) and GPT-5.6 (July 9, 2026). Used as EXHIBITS of self-graded efficiency, not as sources of trustworthy comparison numbers. `[vendor-reported, exhibit]`

 - Grok 4.5 launch page (xAI): shared via https://share.google/8aicinsSQpuKjcNvg - replace with canonical xAI URL at publish. Note the page brands itself "SpaceXAI" while the copyright line reads xAI Corp; confirm the current corporate name before publish.

 - GPT-5.6 launch page (OpenAI): shared via https://share.google/VYe6IfY5qy6xklZth - replace with canonical OpenAI URL at publish.

 - Specific claims cited in the article, both attributed as the vendor's own: xAI's page claims ~4.2x fewer output tokens than a leading rival on SWE-Bench Pro (xAI's own measurement); OpenAI's page leads on tokens-per-task and performance-per-dollar, and its own footnote states the cost and latency figures come from offline simulation and may vary substantially in production.

 - Both pages lean on the SWE-Bench and Terminal-Bench benchmark families, the same public benchmarks Databricks flagged as leaking into training data (see the Databricks entry below). This is the load-bearing link: the industry moved its pitch to efficiency, and the efficiency figures are self-selected benchmarks plus self-run cost simulations.

 - CRITICAL FRAMING: the article does NOT assert any model beats another. It treats every lab identically, Anthropic included, as vendors grading their own efficiency. Do not let an edit turn this into "Model X is more efficient than Model Y," which would repeat exactly the self-graded claims the piece argues against.


## Vendor-commissioned ROI


- Forrester Total Economic Impact 3-year ROI, each commissioned by the named vendor: boost.ai 293%, LogicMonitor Edwin AI 313%, Writer 333%, PolyAI 391%, Cognite 465%. `[vendor-reported]`

 - Writer (333%): https://writer.com/blog/real-economics-enterprise-ai-recap/

 - Cognite (465%): https://www.cognite.com/en/company/newsroom/forrester-total-economic-impact-tm-study-finds-cognite-delivers-465-roi

 - boost.ai (293%): https://boost.ai/guides/forrester-report-the-total-economic-impact-of-boost-ai/

 - PolyAI (391%): https://tei.forrester.com/go/polyAI/PolyAITEI/

 - LogicMonitor (313%): https://www.logicmonitor.com/blog/logicmonitor-edwin-ai-total-economic-impact-study-by-forrester

 - Chart 1 (cover) source. Flagged vendor-commissioned on the chart face. Promoted to cover so the piece leads with the self-grading angle rather than the now-familiar perception-gap image.


- EXCLUDED FABRICATION: a widely reshared "540% ROI across 287 audited enterprise deployments" claim. Traces to a single vendor content page (ajentik.com) with no findable underlying Forrester report. Not used anywhere in the piece; called out in the article as the kind of stat that spreads because nobody checks.

 - Origin of the unverifiable claim: https://www.ajentik.com/insights/enterprise-ai-agent-roi-2026


## Independent measurement


- METR randomized controlled trial: economists predicted 39% faster, ML experts 38% faster, developers' own post-hoc estimate 20% faster; measured result 19% SLOWER. 16 developers, 246 tasks, early-2025 tools (Cursor + Claude 3.5/3.7 Sonnet). INDEPENDENT.

 - arXiv abstract: https://arxiv.org/abs/2507.09089

 - arXiv PDF: https://arxiv.org/pdf/2507.09089

 - Chart 5 source (moved from cover into the independent-measurement section, alongside Faros and NANDA). Strongest single item.


- Faros AI "Acceleration Whiplash," March 2026: 2 years of telemetry, 22,000 developers, 4,000+ teams, comparing each team's low- vs high-AI quarters. Epics/dev +66%, tasks/dev +34%, PR merge +16%; bugs/dev +54%, incidents/PR +242.7%, median review time +441.5%, code churn +861%. INDEPENDENT.

 - ADTmag: https://adtmag.com/articles/2026/04/22/more-code-more-bugs.aspx

 - Matthew Aberham (per-metric breakdown, 2025 vs 2026 bug rate 9%→54%): https://www.matthewaberham.com/blog/faros-acceleration-whiplash

 - Vibe Graveyard (methodology, 50% adoption threshold): https://vibegraveyard.ai/story/faros-ai-acceleration-whiplash-study/

 - Chart 6 source. Note: Forbes rounds churn to "+800%"; the Faros figure is +861%. We use +861% and cite Faros directly, not Forbes.


- MIT Project NANDA, "The GenAI Divide: State of AI in Business 2025": ~5% of enterprise GenAI pilots show measurable P&L impact; funnel for enterprise-grade custom tools 60% evaluated → 20% pilot → 5% live. 300+ deployments, 150 interviews, 350 surveys. INDEPENDENT (academic).

 - Fortune: https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/

 - NANDA report PDF: https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf

 - Legal.io (60/20/5 funnel): https://www.legal.io/blog/5719519/MIT-Report-Finds-95-of-AI-Pilots-Fail-to-Deliver-ROI-Exposing-GenAI-Divide

 - Chart 7 source.


- MIT NANDA "shadow AI" finding: ~90% of employees report using personal AI tools for work (often daily) even where official pilots failed, and this unsanctioned use frequently delivers more value than sanctioned deployments. Same report as above (Fortune + NANDA PDF). Used in the independent-measurement section to isolate the failure to the sold-outcome layer rather than the technology.


## The steelman rebuttal


- Faros bug rate widened 9% (2025 study) → 54% (2026 study) on newer tools. See Aberham (above).

- GitClear: code churn roughly doubled from a pre-2023 baseline (~3.3%) to ~7.1% (2026) across 211M+ changed lines. INDEPENDENT, second source for the rework problem.

 - Larridin (summarizing GitClear's longitudinal analysis): https://larridin.com/developer-productivity-hub/code-churn-ai-era-doubled

 - Chart 8 source. Intermediate years `[interpolated]` between sourced endpoints; endpoints are GitClear's.


## The close / Cisco parallel


- Cisco: real technology, still lost most of its market value in the dot-com repricing. Standard market history; no live-figure dependency.


---


## Stat we deliberately did NOT use (know why)


The Forbes piece cites "the MIT study found AI automation is economically viable in only about 23% of roles." That 23% is real but it is from a 2024 MIT CSAIL paper about **computer-vision tasks specifically** (23% of wages paid for vision tasks), not AI economics in general.

 - MIT CSAIL: https://www.csail.mit.edu/news/rethinking-ais-impact-mit-csail-study-reveals-economic-limits-job-automation

A separate, newer MIT + Oak Ridge study (Nov 2025) puts today's viable automation at 11.7% of the US labor market.

 - Fortune: https://fortune.com/2025/11/27/mit-report-ai-can-already-replace-nearly-12-of-the-us-workforce/

Two different studies, two years apart, two different questions, routinely flattened into one "23% of jobs" factoid. We left it out rather than repeat the misuse. Holding it in reserve as a possible standalone correction piece.


Other Forbes figures available but not used in this piece (would need primary confirmation if added later): Big Tech $740B 2026 capex (+69%); Gartner AI-agent software $207B 2026 (+139%); 115,000+ tech layoffs in 2026; ~$1.3T single-session chip selloff, June 2026; ~95% of enterprise AI usage on frontier models. All `[confirm at publish]`.


---


## Provenance summary


The spine of this piece rests on four **independent** measurements, and that is deliberate: METR (a randomized controlled trial), Faros AI (2 years of telemetry across 22,000 developers), MIT NANDA (300+ deployments), and GitClear (211M+ lines). These are the sharp end of the argument and none were paid for by an AI vendor.


Everything on the other side of the ledger, the resolution rates and the 293-465% ROI figures, is explicitly `[vendor-reported]` and labeled as such on the charts, because the whole thesis is that these numbers are self-graded. That labeling is the point, not a weakness.


The cost-panic figures (Uber, Nvidia, OpenAI) come through a Forbes contributor column and are corroborated by Fast Company and Yahoo Finance, but the exact percentages should be re-confirmed against an exec's own words or a wire report before publish; flagged `[confirm at publish]`. One chart (cost curve) is openly `[interpolated]` and marked ILLUSTRATIVE on its face. One circulating stat (540% / 287 deployments) was identified as an unsourced fabrication and excluded, and one commonly misused stat (the "23% of roles" figure) was left out to avoid repeating an error.


Position/relationship disclosures carried in the piece: Adobe employment (solution consultant) disclosed in the ROI/what-works section; Anthropic/Claude assistance disclosed mid-piece, since Claude is named as a counterparty (Claude Code adoption at Uber). No crypto position is relevant to this piece.


Report Page