Applied AI August 2026 · 15 min read

Frontier or open-weight? Companies don’t choose a model. They choose the level of risk.

TL;DR

  • Frontier and open-weight are not opposite ends of a scale. Frontier describes a level of capability, open-weight a mode of access.
  • Start a new, uncertain or demanding use case on a frontier model via an API, and pick the cheapest level that passes your own tests.
  • Open-weight pays off with steady, high load, a real requirement for local data, or offline operation. An idle GPU is still a GPU you have paid for.
  • The biggest cloud risk isn’t a data leak, it’s dependency. The defense isn’t your own server, it’s owning your tests, data, prompts and a backup model.

In June 2026, Anthropic had to shut off access to its Claude Fable 5 and Mythos 5 models within a matter of hours. It wasn’t a technical outage. The models had been on the market for three days when the U.S. Department of Commerce issued an export directive barring access to them for any foreign national, whether inside the United States or outside it. Because Anthropic had no way to filter users by nationality in real time, it suspended access for every customer at once – including AWS, Google Cloud and Microsoft Foundry. The restrictions were lifted on June 30 and Fable 5 returned a day later, just under three weeks in all.

The lesson here, in my view, is not "the cloud is dangerous, let’s buy our own graphics cards." That would be an expensive and hasty reaction. The real lesson is simpler: if a critical business process rests on a single model, the company doesn’t have an AI strategy. It has a single point of failure.

In the fall of 2025, Google launched its Antigravity development environment. The risk here didn’t show up as a service vanishing overnight, but as its actual economics changing step by step. Early on, the paid plans came with very generous limits that made it possible to run several demanding agentic tasks – tasks where the model takes multiple steps on its own, one after another – with the user practically never hitting quotas in any noticeable way. In December 2025, the daily request count on the free tier dropped from 250 to 20, and in March 2026 Google rebuilt its entire AI subscription: it introduced a system of purchasable AI credits, and on the paid Pro plan the quota reset moved from a five-hour cycle to a weekly one for most models. Developers who were measuring their own consumption reported that before January they could get through more than 300 million input tokens in a week, and then hit a weekly ceiling at just under nine million. These are community estimates, not official figures. What matters is that the plan description stayed almost unchanged throughout.

From a company’s point of view, though, the takeaway is the same: a product that was economically almost unlimited when the flat fee was purchased can, a few months later, turn into a service with substantially lower capacity and additional variable costs. Without any change to your own workflow, contract or technology.

OpenAI has had its own version of the problem. There, the same risk shows up less through an official price change and more through uncertainty about which model the user is actually getting inside a particular product. After GPT-5.5 launched, a number of Codex users reported a noticeable drop in quality on the same kinds of programming tasks. There were also specific reports that the application automatically switched processing to a smaller model despite GPT-5.5 High being selected. That is evidence that an automatic switch to a different, usually smaller model (routing or fallback) can happen; it is not evidence that OpenAI deliberately and systematically passed off GPT-5.4 mini as GPT-5.5, or was quietly trying to save on costs.

Anthropic went through a very similar episode. Claude Code users reported a gradual decline in quality and suspected the company of secretly switching to older models or of trimming the so-called reasoning budget – the time the model spends thinking before it produces its final answer. In April 2026 Anthropic published a postmortem stating that neither the model nor the inference layer had changed. What failed was the product layer: in March, for Opus 4.6, the company lowered the default reasoning level in Claude Code from high to medium in order to shorten response times, a trade-off it later called wrong and reverted. On both Opus 4.6 and 4.7, a bug compounded this by repeatedly erasing the model’s earlier reasoning, which made it seem forgetful and repetitive. Because each change hit a different part of the service at a different time, it felt to users like broad, unpredictable degradation.

For companies, the practical consequence matters more: a model’s name and its settings in the user interface do not necessarily guarantee unchanging performance of the product as a whole. Behavior can be affected by routing, limits, system prompts, tool availability or product updates. So a critical process shouldn’t rely on the feeling that "this model works well," but on regular automated tests that raise a flag when the quality of the service changes.

This is why the frontier-versus-open-weight debate so often starts from the wrong question. Instead of looking for the generally better type of model, it is more useful to ask where we want to pay for convenience, where we need control, and which operational risks we are willing to take on.

A restaurant, a private chef and the recipe

A frontier model can be pictured as a top-tier restaurant. We order the result and the provider takes care of the kitchen, the staff, the ingredients, the hardware, the updates and getting through the lunch rush. We usually get the highest quality available without having to build our own AI infrastructure. At the same time, we accept the operator’s menu, prices and rules.

An open-weight model is more like an experienced private chef we bring into our own kitchen. We get the model’s trained weights and can run it on our own hardware. We have more control over the data, the configuration and when we update the model. The kitchen, the electricity, the security, the servicing and the capacity at peak time, however, are on us. On top of that, we have to reckon with the fact that even a very capable chef may not be able to cook every dish to the same standard. And if twenty guests walk in at once, the chef’s quality isn’t the only thing to solve – we also need a kitchen large enough, and the capacity to serve them.

Open source goes further still. Getting a finished chef isn’t enough. We also need the recipes, the relevant code, information about the training, and the rights to use, study, modify and share the system. In AI these terms often get mixed up, but by the Open Source Initiative’s definition, being able to download the weights does not by itself make something open-source AI.

The important thing is that "frontier" and "open-weight" are not opposite ends of a single scale. Frontier describes a level of capability. Open-weight describes a mode of access. A model can be both highly capable and available for download. Mistral, for instance, describes Medium 3.5 as a frontier-class model and at the same time released it with open weights.

Well-known frontier families include models from OpenAI, Anthropic, Google, xAI and Mistral. The best-known open-weight families include Llama, Qwen, DeepSeek, Gemma and Mistral.

Meanwhile, the gap between top closed and open models is narrowing. New open-weight models often beat older frontier models on specific tasks, in programming or mathematics for example. That still doesn’t mean they are better with ambiguous requirements, long agentic processes or reliable tool use.

This is exactly where companies make their first mistake: they compare two models on a single benchmark and mistake the result for an architecture decision. A public benchmark is not a business process. A model can have an excellent programming score and still fail regularly at classifying one particular company’s contracts.

My default: start a new project on a frontier model

For most new enterprise AI projects, I would start today with a frontier model via an API – a direct technical connection to the provider’s service. Not because the cloud is the ideologically correct answer, but because at the start of a project we know almost nothing that matters for the economics of running it ourselves.

We don’t know how many people will use the solution. We don’t know how long the inputs will be, how many steps a single process will take, or how many errors we’ll find after deployment. Often we don’t even know whether the use case will survive its first three months. Buying hardware at this stage means turning uncertainty into a fixed cost.

An API lets us launch quickly, measure real consumption and find out what level of model we actually need. Very often we don’t need the most expensive model on the market. The right approach is to find the cheapest model that reliably passes our own tests.

That last part matters. Until a company has a representative set of real cases and a clear definition of what a correct result looks like, the debate about models is mostly guesswork.

Cost isn’t decided by the price of a token, but by utilization

Imagine a system that processes incoming documents: invoices, correspondence, contracts or service requests. One document averages 2,000 input and 500 output tokens. The result doesn’t have to arrive immediately; we can queue the documents and work through them gradually.

If a single local server processes a document in five seconds on average, the theoretical ceiling is 17,280 documents a day. That comes to 34.56 million input and 8.64 million output tokens. At OpenAI’s pricing as of this writing, in July 2026, that volume at short-context rates – that is, with shorter inputs – would come to roughly $43 a day (about €37) on GPT-5.6 Luna, or $108 (about €93) on GPT-5.6 Terra. Price lists change, but the point holds: at steady high volume, the API bill starts running into the thousands per month.

An RTX 5090 has 32 GB of memory and a maximum GPU power draw of 575 W. If the whole server pulls roughly 750 to 900 W under load, then at an electricity price of €0.18 per kWh, running it around the clock works out to about €3 to €4 a day, without separate cooling.

At first glance, the payback on your own server looks fantastic. But that comparison only holds if the server is busy most of the time, the local model reaches the quality required, and we don’t need a backup solution, servicing or an extra person to run it.

Change the volume to 1,000 documents a day and the cloud bill drops to a few dollars. At that point, buying a server is much harder to justify. Not because of the price of a local token, but because an idle GPU is still a GPU you have paid for.

Open-weight therefore makes economic sense above all with steady, high and predictable load that we can process on a rolling basis. Typical examples are data extraction, classification, summarization, redaction of sensitive data, or other repetitive document processes.

A company-wide assistant is the opposite problem

A company-wide chatbot or agent looks similar on paper, but economically it is a different product. At two in the morning, hardly anyone is using it. At ten in the morning, ten, fifty or two hundred requests can arrive at once. And the user doesn’t want an answer in four hours; they expect the first response within seconds.

With your own hardware, capacity has to be bought for the peak. That means paying for performance that will sit unused most of the day. Long documents and more concurrent users also consume not just time but memory. A single card that comfortably handles one demanding request may not deliver acceptable response times to ten people at once.

The cloud has a natural advantage with this kind of load. The provider shares a large pool of hardware across many customers, and a company doesn’t have to buy its own cluster just for the Monday morning peak.

So I would start interactive assistants, unpredictable agentic processes and complex tasks with low or uncertain volume on frontier models. I would only consider open-weight once we have real data on usage, response times and quality.

Sensitive data: less panic, more specific questions

Companies sometimes overstate the risk of cloud models, in my view. Serious business and API offerings from the large providers state as standard that customer inputs and outputs are not used to train foundation models without consent. That applies to the OpenAI API, the Claude API, Google Cloud and models served through Azure, for example.

That doesn’t mean the data question can be closed with "they don’t train on it." You have to ask about retention, processing region, subprocessors, logging, connected tools and contractual terms. A public chat application and an enterprise API are not the same product.

Equally, an on-premise solution is not secure in itself – it only moves the responsibility. The company has to handle access control, updates, backups, monitoring, libraries and people on its own.

A local model is a requirement when the data truly must not leave a specific environment, when a contract or regulation demands it, when the system has to work without internet access, or when local processing is part of the product’s value. Not when "the cloud feels dangerous."

The biggest cloud risk is dependency

The Fable 5 case was an extreme but very clean example. The model didn’t disappear because of a technical fault, but because an external rule changed.

So I disagree with the view that vendor lock-in is solved by simply changing an API address. Code can be rewritten fairly quickly. Model behavior can’t. A different model will interpret the prompt differently, use tools differently, refuse different requests and make different mistakes. Real lock-in isn’t only in the code, it’s in the process we tuned to one specific behavior.

The defense doesn’t have to mean your own server. It means owning the tests, the data, the prompts and the business logic. A critical process should have a backup model, and the company should know what happens when the price, the limit or the availability of the service changes.

With an open-weight model we have more control over the version. If a particular model works for us, nobody changes it overnight. That doesn’t mean we have no dependencies. We still depend on hardware, drivers, the inference software (which runs the model), security updates and the people who know the system.

Open weights don’t mean a blank check

Another overlooked topic is licensing and where the model comes from. Open-weight does not automatically mean Apache 2.0 or unlimited commercial use. Some models allow broad deployment; others carry their own restrictions for large users, for redistribution, or for offering the model as a service.

The practical advice for management is this: a model’s license deserves the same scrutiny as the license of any other critical piece of software. You also need to know where the model comes from, what documentation exists for it, and whether it can be trusted in your supply chain.

A European server is not automatically European sovereignty

Many frontier providers now offer European data regions. OpenAI, for example, lets eligible API customers choose to have data processed and stored in Europe. But a server’s physical location is not the same thing as the provider’s jurisdiction. The CLOUD Act, the U.S. law on government access to data, covers data a company holds in its possession or under its control, regardless of where it is physically stored.

At the same time, I wouldn’t read the AI Act, NIS2 or DORA as a general instruction to run models on-premise. Their practical message for management is more like: know the risks, vet your vendors, document your decisions and prepare for an outage of a critical service. DORA, for instance, explicitly requires financial entities to manage third-party ICT risk as part of overall ICT risk management.

European models, then, shouldn’t be picked by passport, nor written off automatically. Mistral is a real alternative, commercially and as open weights. It doesn’t quite reach the absolute top of the American closed models, but the gap today is smaller than most companies think, and for the overwhelming majority of business tasks it is entirely sufficient. European origin is relevant when it reduces a specific legal, operational or strategic risk, not as a substitute for quality.

How we decide in practice

At INSYNAPS we chose an open-weight model for the AI Anonymizer project, for example. The essence of it is that the sensitive document stays on a local server. The model identifies sensitive data and replaces it with a mask or a placeholder value so the modified version can move on to further processing.

Legally speaking, this is pseudonymization rather than true anonymization. Pseudonymized data is still personal data, but the risk of processing it is lower.

Here, staying local isn’t a cost optimization. It is part of the product promise.

For a project handling correspondence with enforcement agents for a health insurer, we went the other way and chose a smaller frontier model. The task is fairly predictable and today’s open-weight models would probably handle it. But the solution was built at a time when the quality gap was wider. It’s a good reminder that the right decision in the original design isn’t necessarily the cheapest one two years later.

For a project with a construction company we stayed on frontier models. The tasks involve ambiguous classification, extraction of the essential information and comparison of requirements. The volume, meanwhile, is relatively low. Buying and running local infrastructure for a small number of demanding cases would be more expensive and probably worse as well.

Three projects, three correct architectures. Which is exactly why it makes no sense to have one company-wide answer for every AI use case.

Hybrid isn’t the goal. It’s the result of measurement

A hybrid approach sounds modern: simple tasks locally, complicated ones on frontier models, and the cloud as insurance against overload. In many companies that will be the right end-state architecture. It shouldn’t be the automatic starting point, though, because two environments, routing and monitoring bring complexity of their own.

I would start with one model, measure quality, volume and errors, and only then split up the individual task types. Once we find that 80 percent of requests are simple repetitive classification, we move that to a cheaper or local model. The complicated cases stay on frontier.

I would be equally cautious about fine-tuning, the additional training of a model after the fact. For most companies it shouldn’t be the first step. First we need good tests, the right context, solid search over internal data and a clear process. Fine-tuning has value on a narrow, stable, high-volume task, for instance when it lets us replace a large model with a smaller one. Training a model before we can reliably measure its errors is just a more expensive way of guessing.

What surprises companies after the pilot

Most often it isn’t the cost, but the number of errors. The pilot gets tested on fifty clean cases; production delivers five thousand ugly ones. Even a one percent error rate means fifty bad cases. A good test set and a human review process are therefore often a more valuable investment than moving to the next generation of model.

The second surprise is the true cost of a single task. With extended-reasoning models, the internal reasoning tokens are billed as well, and the user never sees them in the final answer. On top of that, an agent may call the model several times, use paid tools, correct its own mistakes and re-read the context. So we don’t measure cost by a single answer in a chat, but by the whole completed business process.

Five questions management should be asking

Before deciding, I wouldn’t ask for a list of the best models. I would ask:

  • What quality is actually required, and what does an error cost?
  • Is the load steady, or do we have short, unpredictable peaks?
  • Does the data have to stay local, or is a properly configured enterprise cloud service enough?
  • Do we have the people and the infrastructure to run a model around the clock?
  • What do we do if the model, the price or the terms of the service change?

If you don’t have answers to these questions, it isn’t yet time to buy hardware or sign a long-term cloud contract. It’s time to measure.

A closing recommendation

My practical recommendation is simple: start a new, uncertain or demanding use case on a frontier model. Pick the cheapest level that passes your own tests. Measure the cost of a completed case, the human time and the cost of errors, not just the token count.

Deploy an open-weight model where you have a specific reason: steady, high load, a real requirement for local data, a need for offline operation, existing infrastructure, or a strategic need to keep an unchanging version of the model under your own control.

And even if everything runs in the cloud today, start experimenting with open-weight models before you urgently need them. Not with a large hardware purchase, but with a small test environment and your own data.

Over the next 12 to 24 months, open-weight models will probably keep getting stronger and smaller. The price of a single API token doesn’t have to rise; competition may well push it down. Companies’ total bills can still grow, though, because agentic systems take more steps, use more tools and handle a larger share of the work. So a company’s most valuable asset won’t be the name of the model in its architecture. It will be the ability to find out quickly whether a new model does that company’s particular job better, more safely and more cheaply.

A company shouldn’t buy GPUs because the token bill frightens it. And it shouldn’t stay in the cloud simply because that is comfortable. First it needs to know what a correct result is, what it is worth, and what happens if the service disappears tomorrow. Only then is the choice of model a management decision rather than a technology bet.

Stanislav Gabčo & INSYNAPS

Related

More articles

Applied AI

From Pilot to Production: What Real AI Implementation Looks Like

May 2026 · 6 min

AI Implementation

AI won't fix a broken process. The problem runs deeper than that.

May 2026 · 5 min

Applied AI

70% of enterprise data is unstructured. Classical automation can't touch it.

May 2026 · 5 min

A similar challenge in your organisation?

Tell us about your workflow. We'll tell you whether AI fits – honestly, even if the answer is no.

AI with a sense for your business. Applied AI, not theoretical.

INSYNAPS on LinkedIn

INSYNAPS s. r. o. · Staré Grunty 16, 841 04 Bratislava, Slovakia · ID No. (IČO): 55 819 028 · Tax ID (DIČ): 2122100486 · VAT ID: SK2122100486 · Registered in the Commercial Register of the City Court Bratislava III, Section: Sro, Entry No. 173404/B

© 2026 INSYNAPS. All rights reserved.

No client data ever stays in language models.