Analytics24
3 September
My list

The Long Read · 24 Aug

The middle of the software stack is the casualty

The middle of the software stack is the casualty. The model eats the layer that used to charge rent.

A data-centre aisle in daylight.

Snowflake already gave the 80% line back, and it is winning. Non-GAAP product gross margin sat at 78% in an earlier year and 75.0% on the current full-year guide. Three hundred basis points, in a name that is beating. The 2027 aggregate is already printing.

The series exists. It is called gross margin, and it is moving.

Where the rent sits after tokens cheapen

Model shoptoken list priceWorkflow ownerseats · usagePowerqueue
Schematic of the claim: model shop, workflow owner, power queue. Not a measured series.

The slogan for two years was that general-purpose models make ordinary software cheap to produce. The 2026 numbers are narrower. The casualty is not software. It is the 80% gross margin a generation of SaaS underwriting treated as a law of nature. The middle is still the seat-priced vendor whose product is a workflow a smaller team can approximate with a model and a files folder.

A mid-market software floor at dusk. Seat-based rooms go quiet first.

Two squeezes, one P&L

Application software is being hit from both sides of the margin. Every shipped AI feature adds inference, retrieval, evals, and observability to cost of goods. At the same time the buyer's alternative is improving: a five-person ops team with a frontier model can now build internally what they used to license as seven seats of a vertical tool. Seat-based pricing quietly breaks when one operator with good tooling does the work of five seats. Vendors are caught between a rising input bill and a customer who needs fewer licences.

ICONIQ's 2026 State of AI snapshot put inference alone at roughly 23 cents of every dollar of AI-product revenue at scaling-stage B2B companies, and average AI-product gross margin at 52%, up from 41% in 2024 and 45% in 2025. The trajectory is improving. The floor is also real. Bessemer's 2025 cut had LLM-native companies around 65%. SFAI Labs, looking at mid-market vertical SaaS, describes a 12 to 17 point hit once you add inference (3 to 8% of revenue at scale), eval engineering, observability, and AI-related customer success, landing most honest reporters in the mid-60s to low-70s. Token list prices fell 60 to 75% across 2025. The savings did not fall through, because mature features added retrieval, self-critique, and routing calls faster than the unit price declined.

A $20 million ARR vertical with 80% gross margin has $4 million of COGS, almost all hosting and support. Add an AI feature that is actually used. Inference at 6% of revenue is $1.2 million. Eval and observability take another point or two. Customer-success load rises because the feature fails in ways a dropdown does not. You are now in the high-60s on a good implementation, or the low-60s on a sloppy one, before anyone discusses price. If you also lose 15% of seats because the customer's operator can cover more workflow, the revenue line takes a hit the margin line does not repair. That is the middle. Infrastructure names collect the 6%. The customer collects the seat savings. The vendor collects a press release about attach.

Token list prices can fall while the hall, the memory stack, and the interconnection queue do not. The rent sits below the application layer.

The public names are already confessing, quietly

Snowflake is the cleanest large-cap tell because it always lived closer to consumption economics than seat economics. Product revenue hit $1.33 billion in Q1 FY27, +34%, the strongest sequential dollar growth in the company's history. GAAP product gross margin was 71%; non-GAAP 75%. The company's own multi-year table shows non-GAAP product GM at 78% in an earlier year and 75.0% in the current full-year guide. That is already about 300 basis points, the 2027 aggregate, printing in a name that is winning, not losing. Sridhar Ramaswamy's letter is about Cortex and the Agentic Enterprise. The margin line is about the cost of running that sentence.

ServiceNow is more interesting because it is the seat-and-workflow incumbent. Q2 2026 subscription revenues beat the high end of guidance by 150 basis points, helped by US federal demand pulling some on-prem subscription forward. Full-year GAAP subscription gross margin is guided at 75%, with non-GAAP at 81%. The company said the FY26 gross-margin guide reflects more customers utilizing our hyperscaler partnerships and an acceleration of customer AI adoption. That is a CFO telling you inference is now a COGS line, and that some of it is being paid to Amazon, Microsoft, or Google rather than eaten in a private cluster. The non-GAAP 81% is the number the slide will emphasize. The GAAP 75% is the number that has already moved.

Adobe and Salesforce are less clean, because mix and non-GAAP theatre still dominate the print. The point is that the two names whose economics sit closest to software that runs on someone else's compute and workflow that is growing an AI SKU have already given the 80% line back, in public, while beating. The mid-market names with less pricing power and more seat exposure should print worse, not better. It is the amount a winning data-cloud name has already given back on the non-GAAP product line, and the amount a winning workflow name is guiding toward on GAAP subscription.

Several public SaaS companies began isolating inference-cost ratios in MD&A in Q1 2026, typically 4 to 9% of revenue. Companies that disclose it get analyst credit. Companies that bundle it into infrastructure get the next question. That disclosure split is how you tell a management team that is running the new P&L from one that is hoping the 80% era returns.

Seats were the product

It is a P&L identity. If one operator with an agent does the work of five seats, five seats of revenue become one seat plus a usage SKU. The usage SKU has to carry inference. ICONIQ's 52% AI-product margin is what that SKU looks like before scale and routing discipline. A vendor that keeps the five seats and includes AI is subsidising the heavy user out of the light user's margin. A vendor that splits the SKU has to have a conversation about price that the 2015 playbook did not require. Either way the 80% blended number is the special case, reserved for products that ship almost no AI or that have engineered an unusually cheap call graph.

The historical analog is the 2010s hosting migration, and it is only a partial analog. When software moved from on-prem to cloud, gross margins compressed and then recovered as hyperscaler unit costs fell and vendors stopped over-provisioning. The recovery depended on hosting becoming a small, predictable line. It scales with use, it is paid to a short list of model shops or to a cluster the vendor does not fully own, and the use is the product. A successful AI feature is a COGS event. That is the inversion. In the old world, success was almost-free incremental seats. In this one, success is incremental tokens.

Private equity's old SaaS move was to buy the 80% margin and cut the costs that sit above it. Inference sits below it. You cannot fire a token. Vista's May 2026 note on enterprise inference said the quiet part: in the old P&L the largest cost was people; in the new one the new variable cost is the completion. Caching and distillation are real; they do not restore the underwriting law. The dollars rotate from engineers and support into cost of goods. A roll-up that treats that as a temporary tax will cut the product when it cuts the bill.

Ordinary software, in this piece, has a specific meaning. It is a product whose switching cost was the workflow, not the data, and whose differentiation was a slightly nicer document, ticket, dashboard, or horizontal admin tool. Vertical systems of record that sit on the customer's data, charge for outcomes, and can route models can take a 10 to 15 point haircut and still widen TAM. Feature roll-ups cannot. The middle is the second group. It is also the group whose equity stories were built on net-expansion from more seats in the same account. Net-expansion from more seats is the number that should stall first, while logo growth stays fine and everyone on the call says the customer still loves them. Those two sentences can both be true. The customer loves the product and needs fewer of it.

X is still posting attach. The 10-Q is posting COGS.

SaaS X is celebrating AI attach and agentic workflows as a revenue event. Almost nobody in those threads posts AI COGS as a percent of AI revenue, cost per active AI user, or the gap between top-decile and median user cost. Those are the numbers The SaaS CFO and ICONIQ have been begging finance teams to isolate. Threads that cannot produce them are marketing. The other X feed, infrastructure, is the cross-check. GPU, power, and memory accounts talk about allocation and lead times, not about abundance. If tokens are cheap and clusters are not, the rent sits below the application layer. The PJM piece is the physical version of the same sentence. Electrons are the constraint. Completions are not.

An earlier cut had a slogan that survived because it was true: own the constraint, rent the abundance, the middle does neither. It needed the 10-Q. Snowflake's 71/75 and ServiceNow's 75 GAAP guide are that 10-Q. ICONIQ's 52 is the AI-native version. The cartoon on this page, model shop, workflow owner, power queue, was a diagram of a claim, labelled not a series. The series exists. It is called gross margin, and it is moving.

What this does not say: that infrastructure equities are a free ride, or that every application vendor dies. NVIDIA, the memory names, the power-equipment book, and the data-centre landlords remain a scarcity story with their own lead times (turbines, transformers, HBM). Application software that used to be scarce because of switching costs is becoming abundant because of generation costs. A vertical that uses the abundance to take the customer's workflow, and prices the outcome, can be on the right side of the same shift. A horizontal that used to charge for seats in that workflow cannot. The mid-market basket, seat-exposed, is the print to track through 2027, not the one winner who learned to route to a small model and charge for usage.

Who is actually in the middle

Those two have the data, the workflow, and the ability to route a model and send some of the inference bill to a hyperscaler. They are already taking the haircut and they are beating. The middle is the seat-priced horizontal whose product is a nicer ticket, document, dashboard, or admin console, and the mid-market vertical that never isolated AI COGS and is now including unlimited AI in the seat. Those names will not show up as a single ticker on this page. They will show up as an aggregate: reported gross margin, same accounting basis, through 2027.

Atlassian and Adobe are the historical seat stories people will misuse. Both spent a decade converting a licence to a seat, then a seat to a cloud seat, and both have enough workflow gravity to survive a 10-point haircut. They are not the casualty. The casualty is the vendor whose net-expansion math was more seats in the same account and whose AI SKU is a checkbox. Seat count inside existing accounts, not logo growth, is the tell. Logo growth can stay fine while the customer needs fewer seats of you.

Auditors will force the classification. Inference that is required to deliver the subscribed feature is COGS, not R&D and not opex. Eval-engineering hired to keep the feature from hallucinating in production is closer to COGS than to a research lab. Several Q1 2026 MD&A sections already isolate 4 to 9% of revenue as inference. The names that have not started that sentence are the names whose 80% margin is a timing difference. The original 300 basis-point hypothesis looks conservative against a world where the winning incumbents have already moved that much and the AI-native cohort prints 52%.

What the filings owe

Secondary tells, in order: AI COGS disclosed as its own MD&A line; a shift from seats to usage on the pricing page; and net-new logo growth that stays fine while seat count inside existing accounts stalls. Eval-engineering headcount showing up as a named hire category on calls, SFAI Labs was counting 2 to 5 people at mid-market verticals and 8 to 15 at enterprise platforms in Q1, is a COGS tell that finance will try to park in R&D. Auditors will not let that last forever.

The token wall

OpenAI's flagship output is $20 per million tokens. Anthropic's is $25. Google's is $12. Three official pages, fetched 22 August 2026, already sit on the power theme tile. Those are list prices, not margins, but they set the ceiling on what a seat-priced vendor can include for free. A mid-market SaaS name shipping unlimited AI inside a $150 seat is betting the customer's median call graph stays below a price that fell 60 to 75% last year and will fall again. Snowflake and ServiceNow can route, reprice, and send some of the bill to a hyperscaler. The seat-priced horizontal cannot. That is why the 80% line broke first in the names closest to compute, and why the middle prints next.

The aggregate is the print.