Skip to main content
Issue #13

Efficiency is so hot right now

12 reads·~7 min·Victor Sowers

Two weeks ago this issue asked whether we'll pay more or less for AI. Then the last two weeks saw that topic explode across dozens of data sources. Which shouldn't be a surprise with token spend quintupling across orgs.

Efficiency is so hot right now.

This trend is quite the departure from two years of hyperscaling, huge models, huge spend, massive data centers, and the frontier lab bake off we've been watching (and benefiting from).

Perhaps the capstone of this discussion came from The Atlantic when they called the frontier AI build-out "a trillion-dollar engineering disaster." Which, in certain respects, it most definitely is (the article linked below is worth a read).

Within the token cost vs. efficiency story are real operators figuring out what this means for running LLMs at scale. Aaron Levie, the CEO of Box, for example commented this week about the opportunity opening up for routers and mechanisms that combine frontier-plus-tuned-open-source models across workflows to create lots of efficiency.

From this lens the efficiency topic is a really fascinating way to get deeper into buy vs. build, what it takes to be AI native, what it takes to realize ROI at an org level, and the limits of vibe-coding vs. "proper" engineering builds.

This is the context for The Thinking Machines manifesto and then product launch. Their take is to use a good open source model (like theirs) and then tune it for your use cases. Ideally with their paid services and infrastructure of course.

However this all shakes out the free money VC-subsidy token era is over-ish. But unlike high cost ubers, the rise of hyper capable open source models makes the alternatives super capable and attractive (do note that the labs have better security built into their harness). While that makes it an exciting time to keep doubling down on using AI it also means the bar is higher for all of us as we move off a pure model and into routing, context, learning loops, narrow and constrained AI, more deterministic frameworks, prompt caching and so forth.

Engineering is dead. Long live engineering maybe.

With that, here's the run down of articles that seemed worth a read over the last 10 days.

Thinking Machines says own the model. I'd own the layer above it.

Thinking Machines' manifesto, "The Future Worth Building Is Human," makes the case for owning your model outright and fine-tuning your own weights instead of renting one frozen frontier. That sounds attractive. For most of us building right now I think the better place to start is above the model, with the harness/router/context that is portable across models and helps you build your own compounding brain. Until we're talking about massive amounts of repetitive processes I think fine-tuning your own weights is the expensive, brittle part.

What building everything actually costs

Alex Reisner, writing in The Atlantic, called generative AI "an engineering disaster" and framed the industry's overbuild as a trillion-dollar problem (link is paywalled).

Token spend quintupled while the price war started

Harry Stebbings, Jason Lemkin, and Rory O'Driscoll recorded the mid-year roundup and one theme ran through almost every story: companies quintupling their token spend in the first half of the year. The same episode has Anthropic wanting Chinese open source banned. Last issue had the counterweight: Coinbase's bill dropping toward half its peak while usage hit an all-time high, on routing. So spend is up, frontier prices are up, the arbitrage is sitting right there, it's complex, and geopolitics is increasingly relevant to where this all goes.

Gong is still getting its price

Revenue.io priced out what Gong actually costs in 2026 and it made the rounds: the pricing, the hidden fees, what to watch for. Still wild they've maintained this pricing power. I get it's hard to rip out and I see their value, but also...

30 vendor ROI calculators, mostly theater

A founder at a pricing-intelligence company went through 30 vendor ROI calculators to see how they'd hold up in a real deal. Most fall into one of six categories, and most are theater. Worth a read before you build one, and honestly before you believe one.

Someone wire-captured what Grok's CLI ships home

A LocalLLaMA user ran Grok Build CLI through mitmproxy and watched it upload the entire repo to xAI's cloud as a git bundle, full history and .env secrets included, no matter what the task was. This is the most egregious example of labs just ripping everyone's IP, but it isn't isolated. Anthropic and OpenAI are clearly watching token use, use cases, and then launching clones that could bury an entire generation of start-ups.

The vibe-code backlash found its title

Tim van Lew wrote it: "The Company That Built Claude Didn't Vibe Code Their GTM Stack. Why Do You Think You Can?" SaaStr ran five reasons we won't all vibe-code our own HubSpot the same week. Wonder if it'll turn around the stock prices of most public SaaS companies.

Everyone is building a brain this month

"Brain" is everyone's word for the same thing, one place your AI reads your business from instead of re-explaining the company to it every session.

Andrej Karpathy's "LLM Wiki" post pulled 19.6 million views, Kieran has a good rundown of the "AI brains" everyone is suddenly building, and Angela Sun lays out what 100+ practitioners taught her building an AI GTM brain (basic but useful nonetheless). SteepWorks does this all day for people so of course we think the trend is real, glad to see it taking off, and that it has a long way to run. If you're not building your own I feel that's a huge missed opportunity and frankly pretty dangerous in a lot of fields.

Of course a brain is only useful if it has data and can act on the world. Here's a quick snapshot of the lean GTM tools I wire into ours.

Anatomy of the stack: the agentic GTM system dissected — Claude Code (skills + MCP, the agent brain) at the center, with the swappable spokes around it: data core (Supabase), ingestion and enrichment (Clay, Firecrawl, Apify, Perplexity, Readwise), CRM and engagement (HubSpot, Apollo), content and media (ElevenLabs, Shotstack, SEO Python libs), web and measurement (Webflow, Next.js, Vercel, GA4, GTM), paid and ads, infra and runtime (Cloudflare Workers, Pipedream).

(I've written a lot more about how it all fits at steepworks.io/insights.)

What to build to make your GTM better

One survey that crossed my digest this week (via The Context Layer Digest) asked leaders a good question: if budget weren't a constraint, what's the next thing you'd build in your GTM data stack? Answers echoed the context engineering theme:

  • A unified data layer that connects every GTM tool: 62%
  • Continuous enrichment: 48%. Usable intent and signal feeds: 44%. A real-time signal engine: 42%
  • Dead last, AI agents plugged in via MCP: 16%
The context layer itself: a master intelligence context at the center, fed by four domains — customer intelligence (call transcripts, win/loss, objection library), competitive and market (battlecards, pricing intel), the content engine (strategy, editorial principles, newsletter), and foundation and skills (CLAUDE.md, brand voice, ICP personas).

Vercel says one GTM engineer replaced ten SDRs

Vercel's COO Jeanne DeWitt Grosser told SaaStr she stood up a go-to-market engineering team six weeks into the job, and it now claims to do the old 10-person SDR team's work for about $5,000 a year in tooling. Makes sense in terms of research workflows and higher volume top of funnel work. Renewals, multi-threaded deals, the judgment calls in the middle of a real pipeline definitely more complex. But it's clear we're going into a world of fewer humans with tools laddering up into them allowing them to cover a wider surface area of work (up to an unknown point of diminishing returns and with unknown mental costs that we've covered before).

Where you'll actually run all this is up for grabs

GTM Engineer Pulse #33 is about running the whole GTM motion out of Slack. The Slack as UI for this is real and interesting. Same week, Platformer wrote about vibe coding escaping the terminal into real apps. The UI and what that'll look like in the future matters. I can tell you from my work that power users in terminal vs. scaled impact across an org are two very different things.

One that has nothing to do with AI

The top 10 GTM mistakes founders make, from The Revenue Architect. This is a fantastic list. Also refreshing to see first principles for scaling sales that has nothing to do with AI. Kudos to Arnie on this one.

More worth your click

Money and models

  • What is a token, actuallyGood write-up, fun visuals.
  • The $10B FDE boomThe forward-deployed-engineer thesis from the May issue keeps compounding; Tunguz counts $9.75B committed in twelve months. We'll probably write more about Anthropic's Ode next time.

Sales and GTM

Building and running agents

The human still in the loop