Back to Blog
AI Strategy · 5 min read · Published September 30, 2026

The Case for Putting Open-Weight Models on the Table

Matt Genovese
Matt Genovese
Founder & Product Strategy Lead
A hand reaches in with a screwdriver to adjust the exposed internals of an open server chassis on a workbench, its panel removed and leaning against it. A single cable runs off the desk into a dark, blurred background where nothing is identifiable.

An open-weight model rarely even enters the conversation when a software company picks which AI model to call, because the default locks in before anyone thinks to ask. Here’s the business case for putting one on the table anyway.

When a software company builds AI into a product feature, the choice of which model to call is usually made once and rarely revisited. That’s not because teams can’t reconsider it. Switching means rewiring real infrastructure, not adjusting a setting, so the choice tends to stick. It only gets easier once a team already has the abstraction layer in place to make swapping cheap, which is still the minority case. Anthropic, OpenAI, and Google hold 88% of enterprise LLM API usage between them, and open-weight models actually lost enterprise share this past year, from 19% down to 11%, despite getting more capable. That’s proof open models are rarely even part of the conversation. Here’s the case for putting one on the table.

Figure 1. A single horizontal bar divided into four segments. Three segments, labeled Anthropic, OpenAI, and Google, together fill 88% of the bar’s width. A fourth, smaller segment labeled Everyone else fills the remaining 12%.
Figure 1. Three commercial providers hold 88% of enterprise LLM API usage.

Pricing you can actually choose

Commercial frontier models leave you on metered, per-token pricing set by the provider, regardless of size. An open model opens your options: pay per token, usually less than the frontier rate, or take on infrastructure cost by running it yourself. Planorama’s AI team recently moved an early-stage SaaS client off a commercial API running thousands of dollars a month onto a self-hosted, fine-tuned open-weight model for tens of dollars a month, resulting in a 41x drop in monthly costs.

That math holds at real volume, and the fair comparison is cost per answer, not cost per token. The point isn’t that self-hosting always wins, but that the choice becomes yours.

A Ferrari in the suburbs

A lot of what a feature actually asks a model to do doesn’t need frontier-level reasoning: extraction, reformatting, summarization, basic classification. It’s less the wrong tool for the job and more like buying a Ferrari to drive around the suburbs: the car can do 180 MPH, but speed bumps and school zones mean you’re never going to find out. You’ve paid for capabilities that your tasks won’t use.

By the way, this isn’t all-or-nothing: a common pattern now splits the work, with a frontier model planning the steps and a smaller, right-sized model executing them against the actual data, so the frontier model never touches anything proprietary. And speaking of proprietary data...

The security boundary is yours to keep

Controlling where the model runs is really about controlling what you can promise your own customers. The less data leaves your infrastructure, the more you can commit to when someone asks what happens to it, whether that’s a customer or your own governance, risk, and compliance team. For a B2B company with contractual commitments, that can decide a deal. The risk isn’t hypothetical: in early 2026, a compromised employee at an AI SaaS vendor let a breach cascade through OAuth tokens into a customer’s own systems. Across the big three commercial providers, roughly 3 in 7 identified security incidents since 2021 have involved actual customer data exposure. Security was never the AI provider’s product. It’s yours.

Access and quality you don’t fully control

Getting cut off from your own AI provider is a real liability. A production Gemini API integration was suspended in mid-2026 after hitting a spend cap, with no clear path back in for days. Running multiple providers hedges against this, but you’re still dependent on someone else’s decision to keep serving you.

Content filtering carries a version of this same risk. A provider’s definition of objectionable content is blurry, provider-set, and gets it wrong often enough to be a documented industry pattern, not a one-off bug. That strictness has a cost beyond blocked requests, too: one study found a safety-tuned reasoning model lost 30 points of accuracy on unrelated tasks while cutting harmful outputs from 60% to under 1%. Run an open model yourself, and that tradeoff is yours to set, not inherited from someone else’s definition of what your product is allowed to do. At least that’s the quality risk you can see coming. The next one is harder to spot.

Figure 2. A grouped vertical bar chart with two groups side by side, labeled Before safety alignment and After safety alignment. Each group has two bars: an orange bar for Harmful outputs and a slate bar for Reasoning accuracy. In the Before group both bars are tall, 60.4% and 63.4%. In the After group both bars are short, 0.8% and 32.5%, so the whole chart reads as two tall bars on the left shrinking to two short bars on the right.
Figure 2. One safety-alignment method cut harmful outputs from 60.4% to 0.8%, but reasoning accuracy on unrelated tasks fell from 63.4% to 32.5% in the same test, a 30.9-point average drop.

A dependency that changes without asking

Engineering teams lock down their dependencies to avoid unplanned change. A commercial frontier model is probabilistic to start with, and it can change behind the scenes on the provider’s own timeline. The provider might call that change a general improvement, but there’s no guarantee it makes your specific feature work better. It’s just as capable of making it worse, and you usually won’t know which until it already has. That’s part of who actually owns the proof that an AI feature still meets the bar.

Anthropic’s own postmortem on weeks of Claude Code quality complaints is about as direct an admission as a frontier lab is going to give you: the model changed on its own schedule, and nothing about its name or version number said so.

Self-hosting puts the upgrade path, the validation, and any fine-tuning back in the company’s own hands.

Your own data starts paying you back

A commercial API call is a commodity every competitor has equal access to. Fine-tuning an open model on your own data isn’t. For a larger company, that effort normally pays off in better quality, reliability, and speed, at a lower cost. A smaller company gets those same advantages, and one more on top: a future acquirer will pay for what your fine-tuned model can do that theirs can’t. That’s not hypothetical: a fine-tuned Qwen3-VL-2B model recently beat GPT-5.2 on a CAD-reconstruction task, 72.2% for GPT-5.2 versus 82.1% for the fine-tune. In an M&A conversation, that’s a real asset, not just an efficiency gain.

Figure 3. A small metric title above the chart reads CAD-reconstruction accuracy (Zero-to-CAD test set). Below it, two short vertical bars side by side above a shared baseline. The left bar, labeled GPT-5.2, reaches to 72.2%. The right bar, labeled Fine-tuned Qwen3-VL-2B, is taller, reaching to 82.1%. Both bars are labeled with their exact percentage above the bar.
Figure 3. On the CAD-reconstruction task it was tuned for, GPT-5.2 scored 72.2% versus 82.1% for the fine-tuned Qwen3-VL-2B model.

None of this argues for self-hosting everything. It argues for putting an open model on the list of options actually weighed, against what’s true for your own product and customers, rather than defaulted past because the first commercial provider was the easiest integration. That evaluation, including the migration work behind the case study above, is the kind of hands-on work Planorama’s AI team does alongside the AI strategy work we’ve written about elsewhere.

Matt Genovese
Matt Genovese
Founder & Product Strategy Lead

Matt founded Planorama Design after a career spanning semiconductor engineering and enterprise software, where he saw the same pattern over and over: features that shipped without the requirements and design work that would have made them succeed. He writes about the intersection of AI, product strategy, and the interaction design that carries them.

Related articles

AI Strategy · 3 min read

The Open Model Problem Isn’t the Model

MIT Sloan finds 46% of enterprises stall on open AI models due to integration complexity. The real barrier? Organizations haven’t defined what they need the model to do.

AI Strategy · 3 min read

The Price of Keeping Your Data to Yourself Is Falling

Open models haven’t caught the frontier, but they’re only months behind rather than years, and they run on hardware you control. That changes what keeping your data in-house actually costs.

Let's meet.

Tell us what you're working on. We'll give our honest perspective, and share how we've helped similar teams address their challenges.

Schedule a Discovery Call