> ## Content Index
> Fetch the complete content index at: https://www.taivo.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# Just look at the frontier chart
- URL: https://www.taivo.ai/just-look-at-the-frontier-chart/
- Published: 2026-10-01T06:37:18.000Z
- Updated: 2026-10-01T06:37:18.000Z
- Author: Taivo Pungas
- Tags: Systems, stream

AI keeps getting better. The main driver of this is better LLMs, and they are getting released at a dizzying pace: something notable comes out every month or more often. Keeping track of who now has the best model, by how much, and the various flavors of it, is hard. Even more confusingly, an LLM can be "good" on many axes: capability, cost, speed, personality, etc.

A single chart helps you stay oriented across these developments: intelligence vs cost. Any model can be evaluated like this and becomes one point on this chart.

![Scatter plot with cost on the horizontal axis and intelligence on the vertical axis. Four models: Model 1 is cheap and low, Model 2 is mid cost and high, Model 3 is the most expensive and the highest, and Model 4 costs more than Model 2 but is less intelligent.](https://storage.ghost.io/c/4a/28/4a28c12b-061d-4fef-9be7-3d9f25838589/content/images/2026/10/00-concept-four-models.png)

I like this chart because it captures the central trade-off in picking an LLM at any given point in time: you can get something cheaper and less capable, or powerful and expensive. It's helpful to think of cost as the independent variable: you pick how much you want to spend through your choice of model and reasoning effort, and get the corresponding capabilities.

At any level of cost, one model will be "best" in the sense of "most intelligent". If we only look at that set of models – the best model for any given level of spend – and ignore the rest, we get the Pareto frontier:

![The same four models with a dotted curve labelled "Pareto frontier" through Models 1, 2 and 3, rising steeply and then flattening. Model 4 is faded and sits below the curve.](https://storage.ghost.io/c/4a/28/4a28c12b-061d-4fef-9be7-3d9f25838589/content/images/2026/10/01-concept-cost-vs-intelligence-1.png)

Roughly speaking, you should aim to use a model on or close to the Pareto frontier, whatever your use case. If you don't, you are either leaving money or capability on the table.

So far, we've shown each model as a single point in the cost/intelligence space. That point is mostly not in your control as a user, because it is determined at training time by the model architecture and its training process. But there is one additional knob available to users of most models, both in API and apps: reasoning effort. You can pick a higher effort level to spend more tokens for (hopefully) better results, so different reasoning levels become different points on these axes.

![Models 1 and 3 on the frontier, and Model 2 split into three points labelled low, mid and high reasoning effort. Each higher setting costs more and gains less intelligence.](https://storage.ghost.io/c/4a/28/4a28c12b-061d-4fef-9be7-3d9f25838589/content/images/2026/10/02-one-model-many-points.png)

How does this chart change over time? Labs' R&D work pushes the frontier up and to the left: more capable models, cheaper. Equivalently: one dollar buys you more intelligence today than it did yesterday, and the same intelligence costs less today than it did yesterday.

![Two frontier curves. A faded earlier curve with three faded models sits below and to the right of a darker later curve. Arrows point up and to the left from each earlier model to its newer counterpart: more capable and cheaper.](https://storage.ghost.io/c/4a/28/4a28c12b-061d-4fef-9be7-3d9f25838589/content/images/2026/10/03-frontier-shifts-over-time.png)

We're now ready to look at concrete data! The specific version of this chart I look at is [Intelligence Index vs. Cost per Task](https://artificialanalysis.ai/?ref=taivo.ai#intelligence-comparison-tabs) from Artificial Analysis (AA), a third-party company benchmarking many models, both commercial and open. Their *Intelligence Index* is an aggregated score across tasks from many domains like coding, math, free-form Q&A, and more. *Cost per Task* counts the actual dollars spent to complete the tasks, not just the per-token prices, so it correlates with what you'd see in actual usage: a cheap model that spends many tokens reasoning can end up more expensive. You can read about their methodology [here](https://artificialanalysis.ai/methodology?ref=taivo.ai).

The frontier moves fast, so always refer to AA's website for the freshest data. AA often publishes the results for a model on release day. Here's a snapshot of 01 October 2026, for the frontier plus three main US labs' latest models:

![Intelligence Index versus cost per task (log scale) on 1 October 2026 for Anthropic, OpenAI and Google models, each at several reasoning-effort levels, with a dotted frontier. GPT-6 Luna holds the cheap end, GPT-6.1 Sol the middle, and Claude Opus 5.5 the top at about $6 and 57.6. Fable 5.1, Sonnet 5.5, GPT-6 Astra, Gemini 4 Argon and Gemini 3.8 Flash sit below the frontier.](https://storage.ghost.io/c/4a/28/4a28c12b-061d-4fef-9be7-3d9f25838589/content/images/2026/10/04-frontier-labs-snapshot.png)

Data: Artificial Analysis Intelligence Index v4.3.2, 1 October 2026

Note that cost is usually plotted on a logarithmic scale, which makes the frontier look more linear.

Using model release dates, we see the frontier moving over time. Unfortunately we can't go back much beyond a year, because AA has not comparably evaluated older models, and many have been discontinued by their providers. But even the last 12 months have tons of activity!

0:00 

/0:23 

1× 

Animation: the cost and intelligence frontier moves up and to the left, September 2025 to September 2026\. Data: Artificial Analysis.

Now that you are armed with this concept, beware of its limitations. First, any scoring method will by design lose a lot of nuance, because many domains of capability are compressed into a single number. There can be strong disagreement between perceived quality vs benchmark scores, like Claude Opus 5's insufferable writing style regardless of its high benchmark scores.

Second, you might not care about being on the frontier. There are many other factors than cost and intelligence that go into a model choice: intelligence on your specific tasks, speed, personality, licensing, vendor availability, SLAs, developer experience, geopolitics, switching cost, etc. The easiest way to find an appropriate one for yourself is to shortlist a few you could use, and then run an evaluation on the tasks you care about.

And third, for some tasks, you might find a specialized solution wildly beats the LLM frontier! For example, language detection using [trigram statistics](https://github.com/greyblake/whatlang-rs?ref=taivo.ai#how-does-it-work), speech-to-text using a [model running on Apple Silicon](https://desertant.com/models/voz/?ref=taivo.ai), or even having AI [train a specialized ML model](https://huggingface.co/chat/?ref=taivo.ai) specifically for your use case.

Next time you feel overwhelmed by everything that's released, keep your sanity by anchoring in the cost-vs-intelligence chart.