You’re Evaluating Frontier Models Wrong

The Closeness Penalty

You are evaluating Frontier Models Wrong. You should shave off at least 10-20% for the very fact they are closed.

Small note: at the time of writing I equate frontier with closed proprietary, as opposed to open-weight. We may be blessed by open frontier later but today we sadly have to consider compromises.

Take a thousand tasks and run them across open and closed models. The closed model’s terms of service will block some of them. Refusal, out-of-service, content filter triggered – whatever the reason, that task scores zero.
If you’re evaluating each task as pass/fail, those zeros add up fast. The closed models score lower not because they’re less capable, but because they’re limited by the companies who provide the service.

Which should be accounted for in your own personal benchmark for your usage.

Not even to say that these companies may dislike you in particular.

If you have the wrong nationality, or you are in the wrong country, or they suddenly did not like how your IP looks, or they think that your payment is wrong for some reason.

Also think of benchmarks on the long scale. If you run it for a whole month of work and half of it is downtime just because they restrict your access for some reason, then you can shave 50% off of it.

So I invite you to think not about the fancy graphics you see on a benchmark and go "oh, this is good." You should think about the total lifetime of usage.

Your Workflow Is Hostage

With open models, you set up your infrastructure once and you can be sure about it. Yes, they may be dumber, they may be worse – but you build a pipeline that works, and it won’t be broken. Not by price hikes, not by filter changes, not by some provider deciding to ban you overnight.

Sure, frontier models might deliver greater net benefit by being more capable on some tasks. But how many of your workflows actually need the bleeding edge? And what’s the real cost of ownership across, say, six months – not just the API bill, but the cost of your pipeline breaking when the provider changes something?

The dependency is itself a fine you carry on top of the bill.
This is not a tool one department uses and can painfully switch out. It’s something most businesses jump into – one provider, one model – and then adapt their whole workflow to it. Which is precisely why any switch can be devastating, the moment the provider withdraws something or ships an update that makes the work impossible.

There’s an assumption that models will get smarter and your current workflows will become outdated anyway. But the reverse is just as likely: your workflows break in the next release because they cranked up censorship, or the model got too clever for its own good and stopped being adaptable to your setup. With open models, you can always be sure you have the original. With closed, you can never do that.

And yes, you’re expected to keep switching anyway, because every month something new shows up. But the switching is itself the risk. The new model gets smarter – and destroys the workflow you built around the old one. When the option you depend on can go unavailable or degrade after the fourth update, treat that as a risk, not an opportunity.

Remember picking your CRM? That thing is the nervous system of your business. Imagine being forced to change it tomorrow – the pain, the migration errors, the downtime. The processes you set in place need stability. They can’t drift, degrade, or get yanked away.

You can only guarantee that when you control the software yourself. And we haven’t even mentioned critical applications – military, medicine, energy – where the cost of a surprise breakage is catastrophic.

How can this break? Let’s count:

  • They degrade the model. Make it worse. Accident or intention, doesn’t matter.
  • They change the sampling parameters or inference settings, and your performance tanks.
  • They rewrite the terms and conditions whenever they feel like it.

Proprietary vendors – by intention or error – can become your enemy overnight.
And that’s before any fucking government leans on the same companies, adding a whole extra layer of danger to your business.

Remember how everyone screams that it’s a bubble and all that blah blah blah, those companies are gonna collapse and all that?
Take this as a small possibility. Do you want to go down with them because you relied on their service always being there?

You’re Training Your Replacement

The game is not win-win.

On paper it looks nice. You pay a company for a service that makes your business more efficient, giving you an edge. But here’s what actually happens: they use the telemetry. The data you send. They train the next generation of models on your business data, your expertise – the stuff your engineers put into prompts, intentionally or not. They save it.

This is straight up corporate espionage without making it visible.

Some people are fine with this. Share data, models get better, models serve your business better. You might bet that you’re moving fast enough to keep an edge even as they improve. But each time you use them, they get more powerful. To stay ahead, you not only have to push harder – you have to hope they don’t dump you later. To me, that’s worse than a bad deal.

Others see it as an existential threat. The provider’s growing capability doesn’t just let them replace you faster – it lets them hand the same knowledge to your competition. And you might not adapt quickly enough to stay ahead.

So it’s a double-edged sword. Some businesses voluntarily share data because they believe they can accelerate fast enough to keep an edge in their niche. Others see it as a direct threat to their existence. One thing is certain: if you keep everything on-premise and never send the data, they can’t extract your specific know-how, your business intelligence, the edge you’ve built.

But the proprietary nature of the company means they get to choose. They decide whether to serve you, your competition, or someone else entirely. They can pick who becomes your competitor, armed with the same tools, trained on the insights from your own business.

The ideal world looks different. Everything open. Models trained on a shared corpus of domain knowledge donated by businesses. Each company keeps its edge by being fast enough to implement and use whatever hasn’t made it into the datasets yet.

This ideal breaks on two realities.
First, the same company trying to make money is also building the tool to replace you – not just providing it for everyone. Second, that same company gets to choose who it does business with.

It’s one of the rare cases where a company can use you as a client to improve the tool, and then use that tool to replace you entirely.

Real Price vs. Sticker Price

Even if you set aside the closeness penalty and the reliability risks, there’s the price.
Which is always higher than it looks.

If you take an open-weight model and host it yourself, you pay inference cost.
For a closed API, you pay inference cost plus the provider’s markup – and that markup isn’t just profit. They’re pricing in future R&D, the next model, the one after that. You’re paying a development tax on top of the service.

Worth it for the task? Maybe now, maybe not tomorrow.

And the price is static. Five bucks per output, forever – or until they release the next version and reconsider.
Or even worse – change their terms of service and suddenly you land in an "extra tier" or something.

Potential to optimize.

Meanwhile, with open models, the cost curve points down. New optimization methods arrive. Competition between hosting providers drives prices lower. You yourself can quantize, adapt hardware, and squeeze out another 10% in savings. You’re not stuck with whatever number someone in a pricing meeting picked last quarter.

When you’re evaluating model fit for your business, here’s what actually matters: can you rely on it long-term? If one provider dies, you switch to another. You pick the one that matches your caching strategy, your rate limits, your hardware. With closed, they can pull the plug whenever – shut it down, ban you, change the terms. Your business goes with it.

You can change the model. Modify it, add LoRAs, train it on your own data. You can improve it for your specific domain by 10-20% – which you simply cannot do when it’s closed. Fine-tuning on your use case is a capability that doesn’t exist in the API world.
Many industries have that data they are even unable to use with anything proprietary – trade secrets or some sensitive/private – which may never even be touched and the workflows may never be automated if you have to have to use only the cloud solution.

And you control your own spending. Access to weights means you can quantize, adapt to your hardware, lower costs, optimize.

Even the current providers are very inflexible. They give you one choice and that’s all. Few options for a premium tier, a slower tier for batch tasks, or any flexible arrangement at all.
You’re forced to pay whatever the provider demands.
The model shifts from being a rented service to being a tool you own.

Many would still prefer to host it in a data center, but once you scale big enough, things start to matter.
It is already the norm as your business grows, as your infrastructure demands more, to rewrite the software from scripting languages to compiled ones – to change from Python to Rust, for example.

The models are just another software component. Why should we treat them any differently?

The Missing Harness

We’re stuck in a weird paradigm. On one side: the static model, weights frozen, unchanging. On the other: the flexible harness – the code you wrap around it. But we’re missing a third layer: changing the runtime of the model itself.

We prefer to keep the weights frozen at the checkpoint. That part stays.

The missing harness is the full range of model runtime parameters. Right now, the tools are primitive – adjust sampling, tweak temperature, the basic stuff the inference engine exposes. But this is an incredibly narrow view, and we haven’t even discovered most of what’s possible.

Think of the machine harness as an autonomous or semi-autonomous layer that sits between the model and your code. Without it, you’re leaving a ton of customization on the table that could give you an edge over the base model.

What could it do? For starters: greater stability when generating code or running critical applications. Better control over instruction following through methods we haven’t built yet. Fine-grained control over creativity – not just a stupid temperature number, not just banning words from the logits. You could ban and enhance concepts. Activate steering toward a specific style. Force behaviors.

This sounds complicated. But mostly because nobody has paid attention and nobody has developed it enough to be useful. We’re already using coding harnesses – mostly vendor-supplied – and plenty of us tinker with them enough to customize their behavior. We already use vibe coding to improve the harness itself. We could do the same for model runtime: express what we want in natural language, run evaluations inline, confirm the output, iterate.

If that intervention gives you even 5% more performance, then any open model is 5% better than any closed one on the same benchmark. That’s performance you’re leaving on the table.

And here’s the security dimension: only this harness lets you access the model’s hidden thinking. You can analyze it, categorize it, know when it’s scheming, betraying you, or doing just stupid shit.

Do you want to trust someone else to do this thing? Are they even doing it? Are they even doing it by your rules?
Whatever important thing you’re doing, you’re better off when you control the inference parameters.

Conclusion

Benchmarks measure capability in a vacuum. They don’t measure whether the model will still serve you next month, whether it’ll refuse the task you actually need, or whether your data is being harvested to build your competition.

The real evaluation accounts for longer time scales and whether you are able to improve and optimize in the future.

If you are still picking, today prefer to take something open.
If you are already knee deep with corporate contracts and so on, consider the gradual phasing out of anything proprietary before it catches up with you.


upd 24.09: added a small note about frontier === proprietary