Is Rapid AI Price Deflation A Gift Or An Existential Threat To SaaS?

Building with OpenAI models just became dramatically cheaper overnight. The company reduced GPT-5.6 Luna API costs by 80%, lowered GPT-5.6 Terra by 20% and revealed a Fast mode for GPT-5.6 Sol that accelerates output speeds up to two and a half times. OpenAI credits improved infrastructure efficiency. The cuts took effect immediately across all accounts.

Software founders who mapped out business models when AI inference was expensive are seeing more than a routine price cut. This shift forces them to rethink the core financial assumptions behind their startup. The intelligence layer that cost a dollar per million tokens in January costs twenty cents today. Competitors who previously struggled to fund AI capabilities at scale can suddenly afford to do so, and unique selling points from six months ago are quickly losing their strength.

 

The Gift Case

 

The optimistic view makes a strong case. Lower AI inference costs expand what becomes commercially viable. Always-on agents, deep personalisation, real-time analysis and processing complete customer histories were possible before, but the economics made little sense. These capabilities are now cheap enough to build into every application. For software teams holding back on AI features because the margins were too tight, an 80% price drop is an open invitation to build.

Gross margins are another clear advantage. Now that existing AI features cost less to run, unadjusted software pricing directly improves your bottom line. Reports indicate that lower prices will spur faster enterprise AI rollouts instead of cutting overall IT spend. Higher usage volume offsets lower per-token revenue for providers, and that same logic means SaaS companies can use lower serving costs to fuel faster product development and distribution.

 

The Crisis Case

 

A tougher look at the market shows that commoditised AI inference erases established strategic edges. If your advantage relied on offering features competitors couldn’t afford, that advantage is gone. High costs no longer stop smaller rivals from offering AI features at scale. Every competitor now gets access to the same AI tools for the same price.

This hits hardest for companies that sold themselves as AI-first, raised money by claiming their AI setup was proprietary or too expensive to copy and built their pitch around early adoption. Those companies now face tougher questions. If your core intelligence is delivered by an API that anyone can call for twenty cents per million tokens, what exactly are customers paying you for? The answer is likely workflow depth, proprietary data, integration lock-in, domain expertise, distribution or brand. If it’s primarily the AI itself, the conversation with investors is about to get harder.

We asked SaaS founders, operators and investors to weigh in on what the price cut actually means for their businesses, their pricing and their differentiation story.

 

Our Experts

 

 

  • Ashish Nagar, CEO and Co-Founder, Level AI
  • Pierre Rogers, Co-Founder and Director of AI Video, FuguTech
  • Patrick Gibbs, Founder, Epiphany Dynamics
  • Yerbol Baimencheyev, Founder and CEO, HeadhuntEZ
  • Sandip Patel, Senior Cloud Solution Architect, Microsoft
  • Tony Wenzel, Founder, Excipio
  • Michał Piszczek, CTO, Archdesk
  • Viktor Bulanek, Founder, Penetrify
  • Luke Budka, AI Director, Definition AI
  • Mykyta Chernenko, Founder, AIWriteBook

 

 

Ashish Nagar, CEO and Co-Founder, Level AI

 

 
Ashish Nagar, CEO and Co-Founder, Level AI
 

“This price cut is the best thing to happen to serious AI companies, and a reckoning for everyone else. I’ve said for a while that per-token costs had to fall five to ten times for the AI app economy to actually work. OpenAI just did most of that in a single day, three weeks after shipping the model. That isn’t a shock. It’s the intelligence layer becoming what it was always going to become: cheap, abundant, and available to your competitor as a one-line API change.

“So the real question lands hard. If your pitch to investors was ‘we adopted expensive AI first,’ you never had a moat. You had a head start, and it just evaporated. The founders panicking this week are the ones who wrapped someone else’s model and called it a product. It changed nothing about how Level AI prices or how I talk to investors, because we never sold access to a model. We own our model layer for customer service, train it on our customers’ data, and it improves only for them.”

 

Pierre Rogers, Co-Founder and Director of AI Video, FuguTech

 

 
Pierre Rogers, Co-Founder and Director of AI Video, FuguTech
 

“OpenAI’s 80% cut didn’t commoditise differentiation. It revealed that most ‘AI-first’ SaaS never had any. A large share of the category was a 10x markup on someone else’s API with a nice onboarding flow. If your moat evaporated on 30 July, it was weather, not a moat.

“It hasn’t changed our value story, because we never told customers that access to a model was the product. The story that survives inference deflation is built on the things that get more valuable as intelligence gets cheaper: proprietary data, distribution and workflow depth. Cheap tokens mean we can personalise at a scale that was cost-prohibitive six months ago. That’s margin and capability, not a crisis. Stop selling intelligence. Sell what you know that the model doesn’t.”

 

Patrick Gibbs, Founder, Epiphany Dynamics

 

 
Patrick Gibbs, Founder, Epiphany Dynamics
 

“I build and price agentic workflows for small businesses: intake bots, scheduling agents, lead routing. My pricing has always assumed inference cost was a real input, not free. When OpenAI drops GPT-5.6 Luna’s API cost 80% overnight, it changes what I can promise a client for the same monthly fee, and it changes what a client’s own team could plausibly stitch together themselves with a cheap subscription and an afternoon.

“The differentiation question isn’t ‘can I access the model’ – everyone can now. It’s whether I understand the specific business well enough to know which three decisions actually need automating and which twelve don’t. That judgement doesn’t get cheaper when the API does. If anything, cheaper inference makes more businesses try to DIY the easy 80% and come looking for help on the hard 20%, which is where the margin already was. I’m passing the savings forward as capability, not as a lower price. More automation for the same retainer, not a discount.”

 

Yerbol Baimencheyev, Founder and CEO, HeadhuntEZ

 

 
Yerbol Baimencheyev, Founder and CEO, HeadhuntEZ
 

“The price cut was good news for us and a little embarrassing for the industry, because it exposed how many pitch decks were quietly built on the API bill. If your moat was that intelligence used to be expensive, you never had a moat. You had a head start on a subscription.

“We made a decision early that saved us from this moment: we never priced the AI. HeadhuntEZ is an AI headhunter – the system runs the whole candidate search, but clients pay per booked meeting and a success fee only when they hire. Nobody has ever paid us for tokens, so tokens getting 80% cheaper doesn’t touch our value story. It just improves our margins and lets us profitably serve smaller companies that couldn’t justify a traditional recruiting fee. Any competitor can call the same API tomorrow. What they can’t drop in is eight years of recruiting judgement encoded into how the system searches, and a human quality bar on top.”

 

Sandip Patel, Senior Cloud Solution Architect, Microsoft

 

 
Sandip Patel, Senior Cloud Solution Architect, Microsoft
 

“The real headline is not that GPT has become cheaper. It is that intelligence is rapidly becoming abundant. Lower inference costs are good for the software industry because they remove a constraint on what AI can do at enterprise scale. But buyers aren’t investing in tokens. They’re investing in software that fits how their organisation actually operates, can be governed and delivers measurable outcomes.

“When AI becomes cheaper, the value doesn’t disappear – it moves. The winners won’t be the companies with the lowest model costs. They will be the companies that own the customer workflow, the proprietary data and the business outcomes. The conversation with investors is changing. Instead of ‘what model are you using?’, the more important question becomes ‘what unique data, integrations and domain expertise do you control?’ AI capability is becoming commoditised, but customer context is not.”

 

Tony Wenzel, Founder, Excipio

 

 
Michał Piszczek, Chief Technology Officer, Archdesk
 

“OpenAI just made a strong argument for us. A capability that somebody else can reprice 80% overnight was never a moat. Founders who sold AI-first as the differentiator were selling a rental agreement and calling it real estate.

“Our pricing didn’t change and neither did our investor story, because neither was built on privileged access to expensive models. What the cut doesn’t fix is the shape of the spend. Datadog’s 2026 State of AI Engineering report found that 69% of enterprise LLM input tokens are system prompts and repeated context. Companies are paying, over and over, for answers they already own. Cheaper tokens don’t reduce that behaviour. They subsidise it. Price per token fell 80% while agentic call volume keeps climbing, so anyone modelling this as an 80% cut to their inference line is watching the numerator and ignoring the denominator.”

 

Michał Piszczek, Chief Technology Officer, Archdesk

 

 
Michał Piszczek, Chief Technology Officer, Archdesk
 

“The 80% cut did not change our pricing. It changed which number we defend in front of investors. When we priced our AI features, the model bill was never the expensive part. Run the numbers: a fully loaded agent-hour, compute plus the human review minutes it triggers, comes out around 0.38 times the human hour it replaces, and the review minutes dominate that figure, not the tokens. Cut inference 80% and you move the smaller term.

“What does not commoditise is everything around the model: the domain data, the workflow it sits inside, and the evidence that the output is correct. In our agentic pipeline, verification is 80 to 90% of the real cost. A developer agent burns $2 to $10 of inference per task. The QA agents that prove its work cost $30 to $150. Routing the routable 60% of a workload one tier down cuts the bill about 48%, and we spend that difference on verification capacity, because that is what caps how much AI work an organisation can absorb.”

 

Viktor Bulanek, Founder, Penetrify

 

 
Viktor Bulanek, Founder, Penetrify
 

“We don’t run on OpenAI – we run on Claude – so the 80% cut didn’t hit our bill directly. It still matters, because it sets the price expectation for everyone. Our honest position is uncomfortable to say out loud: the model was never the moat, and we learned that by accident rather than by strategy. We’ve swapped the underlying model twice in a year. If a component can be replaced in an afternoon, it’s an input, not a differentiator.

“What is actually hard is everything wrapped around the intelligence. Running offensive tooling safely. Checkpointing a four-hour autonomous scan so an interruption doesn’t destroy it. Throttling parallel agents so they don’t trip provider rate limits. Producing evidence an auditor will accept. That took months, and none of it got 80% cheaper on 30 July.

“The part I think founders are missing: cheap inference arms attackers too. Coldcard’s maker suspects an attacker used AI to find a five-year-old bug that its own automated review had missed weeks earlier. Eighty-nine million dollars left in forty-one minutes. For a security vendor, token deflation grows my market faster than it commoditises my product.”

 

Luke Budka, AI Director, Definition AI

 

 
Luke Budka, AI Director, Definition AI
 

“We built a model-agnostic AI platform. Our moat is human expertise. If our creatives tell us model A is great at tone of voice, model B is outstanding at photorealistic imagery and model C is the one for realistic video dialogue, in they go. The token price economics don’t matter to us; we pass on the cost savings. What our clients are paying for is our expert evaluation and benchmarking – the AI field is so fragmented that clients have no way of working it out themselves, particularly because frontier labs don’t focus on benchmarks that matter to marketers, and public benchmarks are based on votes from the public, not creative experts. Any model-plus-prompt wrapper business is already struggling, so I can’t imagine API price fluctuations will make much difference to their longer-term sustainability.”

 

Mykyta Chernenko, Founder, AIWriteBook

 

 
Mykyta Chernenko, Founder, AIWriteBook
 

“The cut changed nothing about our pricing, because the model was never what customers were paying for. If the intelligence layer were the product, an 80% price drop would have gutted us, because any customer can call the same API themselves for pennies. They don’t, and the reason is the useful part.

“More than 11,000 writers have put over 190 million words through us and about 4,000 book-length manuscripts came out the other side. Fewer than around 90 of those have an external buy link on them. The expensive part of this business was never generating text. It’s everything wrapped around the distance between a finished draft and a book somebody can actually buy: holding a story consistent across a whole book, translation, cover, export, the publishing steps. None of that got cheaper on 30 July.

“My honest read on the founders in trouble is less comfortable. If an 80% input cost cut is an existential event for you, what you were selling was access rather than a product. That was always temporary. The cut just set the date.”