Cost Per Token Is the Wrong Metric
By Tony Ojeda
There is a common argument in AI that open-weight models are cheaper than proprietary models from companies like OpenAI and Anthropic. The logic seems straightforward: if the weights are free, you avoid paying a provider every time you send a request. But the weights are only one part of the cost.
Running the model still requires compute, usually on expensive GPUs. Those GPUs have to live somewhere, someone has to operate them, and the model still has to perform well enough to justify using it. The real decision combines model capability, deployment method, cost, and control.
Free Weights Don’t Mean Free Inference
When a company releases an open-weight model, it gives you access to the model parameters. It does not give you the infrastructure required to run them. If you deploy the model yourself, you need GPU capacity. That could mean buying hardware, renting GPUs in the cloud, or paying for dedicated inference infrastructure. You also take on some combination of deployment, scaling, monitoring, networking, and reliability.
Proprietary APIs hide all of that. You send OpenAI or Anthropic a request, they run the model, and you pay based largely on usage. That creates two different cost structures. With an API, cost rises with the number of tokens you consume. With self-hosting, much more of the cost is tied to the infrastructure you provision.
This is why utilization matters. If you only make occasional requests, paying by the token is usually attractive because you pay nothing while the model is idle. If you have a large, predictable workload and can keep GPUs busy, self-hosting can become cheaper because the incremental cost of another request falls substantially. What an open model gives you is the option to trade per-token pricing for infrastructure cost.
You Don’t Have to Host the Model Yourself
There is also a middle ground. Providers like Together AI, Fireworks, and Groq host open models and expose them through APIs. From a developer’s perspective, the experience looks much like calling OpenAI or Anthropic: send a request, receive a response, and pay for what you use. This can give you much of the cost advantage of an open model without requiring you to operate the infrastructure yourself. That leaves at least three options:
Proprietary API → hosted open-model API → self-hosted open model
A proprietary API gives you access to some of the strongest models with very little operational burden. A hosted open-model API can provide much cheaper inference when the model is capable enough. Self-hosting offers the most control and can produce the lowest marginal cost when usage becomes large and predictable.
Measure Cost Per Successful Task
The problem with comparing prices directly is that the models are not equally capable. Suppose an open model costs one-tenth as much as a frontier model (one of the most capable models available today) but succeeds on only 75% of your tasks. If failures are cheap, that may still be a good trade. If every failure requires an engineer to spend twenty minutes fixing the result, the inference savings disappear quickly.
The more useful metric is cost per successful unit of work, which includes inference, retries, validation, human correction, and the consequences of bad outputs. For classification, tagging, structured extraction, routing, or basic transformations, an inexpensive open model may perform well enough that paying for frontier-level reasoning adds little value. The calculation changes for difficult reasoning, coding across a large repository, long-running agents, tool use, research synthesis, or situations where the model has to recover from unexpected conditions. In those cases, a more expensive model may still be cheaper overall because it completes more of the work correctly. The right comparison is which model produces acceptable results at the lowest total cost. I think about my own systems this way. The autonomous remediation system I built for data pipelines tracks what each investigation, fix, and revision costs, and caps spending per run, because what matters is what one resolved issue costs, not what a token costs.
Different Tasks Need Different Models
Once you think about the problem this way, using the same model for an entire AI system starts to look inefficient. Imagine a research system that collects documents, determines relevance, extracts structured information, summarizes evidence, resolves contradictions, and produces a final analysis. Those steps do not require the same level of intelligence.
A cheap open model might handle document classification and extraction perfectly well. A stronger model could deal with harder summarization or ambiguous evidence. A frontier model might only be necessary for final synthesis or unusually difficult cases. The system can escalate work based on difficulty rather than paying the highest intelligence cost on every request.
That is not particularly different from the way we design other computing systems. We do not use the most powerful hardware for every operation. We allocate expensive resources where they produce enough additional value to justify the cost. AI systems should work the same way. I made a related argument in The Harness Changes the Model: the useful question is rarely which model is best, but which combination of task, model, and environment fits the job.
Security Is a Separate Dimension
Open models also provide deployment options that proprietary models do not. You can run an open model on hardware inside your own data center, where prompts, documents, and outputs never need to leave your network. Self-hosting can also mean deploying the model inside your own AWS, Azure, or Google Cloud environment. The cloud provider owns the hardware, but you control much more of the surrounding environment: networking, storage, access policies, encryption, and logging.
This creates a spectrum of deployment options, from external APIs to managed private services, private-cloud deployments, on-premises infrastructure, and fully air-gapped systems. The important distinction is who controls the execution environment and where the data travels. An open model running through a third-party API still sends your data to another company. A self-hosted model inside your own environment gives you much more control. But a poorly managed private deployment can still be less secure than a well-run cloud service. Open weights provide the option to choose that boundary.
The Real Advantage Is Flexibility
That flexibility may be the most important advantage of open models. You can start with a serverless API when usage is low. As traffic grows, you can move to dedicated capacity. If economics, security, or compliance justify it, you can later deploy the same model inside your own cloud environment or on your own hardware. Proprietary models offer a different advantage: access to capabilities you may not be able to match with an open model, without having to manage the underlying infrastructure.
Those choices do not have to apply to an entire system. A production application might use a cheap hosted open model for routine work, a frontier model for difficult reasoning, and a privately deployed model for sensitive data. That is probably where this is heading: systems that use different models for different jobs.
Ask What the Work Requires
Whether open models are cheaper than OpenAI or Anthropic depends on volume, utilization, model quality, operational overhead, and the cost of failure. The questions that settle it are practical. Which model is capable enough to do the job reliably? What is the least expensive way to serve it at the expected volume? How expensive are failures and human corrections? And how much control over the data and infrastructure does the workload require?
Once you answer those questions, the choice becomes much less ideological. Use the cheapest model that reliably does the job. Pay for more intelligence when it materially improves the result. Take on infrastructure when the economics or security requirements justify it. The future is probably AI systems that are more deliberate about where each kind of intelligence belongs, open and proprietary alike.