No provider influences the numbers
No provider pays for or influences the barometer, and none sees results before they are published. Aprex itself uses models from several of the providers in the solutions we build.
Tasks resemble what employees do with documents, emails and minutes, and each has an answer that can be checked.
Answer keys and scoring rules are written before any model answers, and they do not change after we see the results.
Same instructions, same settings and same service for every model.
Most tasks are never published, so they cannot end up in training data and give some models a head start.
Scores come from fixed rules wherever possible. Judge models are used only where quality must be rated, and count for at most 40 percent.
48 tasks show direction, not decimals. The confidence interval shows how certain the differences are.
Score, cost and response time sit side by side and are never mixed.
No run is published before a person has reviewed the outliers.
Each of the six types has eight tasks. Two per type become public with their answer key, so anyone can see what the models get. Six are published now, and the rest follow when the task set is complete. The other 36 stay hidden. The tasks are in Norwegian.
| Types | Document extraction, Customer replies, Summaries, Regulation, Classification, Nynorsk |
|---|---|
| Content | Fictional companies, people and documents, or public sources. No customer data, not even anonymised. |
| Language | The Nynorsk type is written in Nynorsk. At least one task in each other type has Nynorsk content. |
| Regulation | Answer keys are checked against the statute on Lovdata before each new version of the task set. |
| Each task has | Task text, answer key, a correct answer in other words, a deliberately wrong answer and scoring rules. |
| Check | The key and the reworded answer must score at least 95, and the wrong answer at most 40, before a task can be used. |
| Service | Vercel AI Gateway, one shared access point to every provider. The exact model ID is logged. |
|---|---|
| Instructions | «Du løser en arbeidsoppgave for en norsk bedrift. Følg instruksjonene i oppgaven nøyaktig, og svar bare med det oppgaven ber om.» (Norwegian: you solve a work task for a Norwegian business; follow the task exactly and answer only what it asks.) |
| Temperature | 0 where the model supports it |
| Reasoning | Provider default. The level is logged. |
| Tools and web search | Off |
| Max answer length | 2000 tokens |
| Time limit | 120 seconds |
| Repeats | 3 per task |
| Errors | On network errors or overload we retry once. If it fails again, the answer scores 0 and counts as an error. |
| Hidden tasks | Sent only to providers that do not train models on the data. |
| Task score | Mean of the three runs |
|---|---|
| Type score | Mean of the eight tasks in the type |
| Total score | Mean of the six types. All types count equally. |
| Confidence interval | We resample the tasks with replacement 10,000 times and compute the total each time. The interval is the middle 95 percent. Models with overlapping intervals count as equally good. |
| Stability | Spread between the three runs |
| Cost | Actual price per call, including hidden reasoning, scaled to 100 tasks. Given in US dollars excluding VAT, because providers bill in dollars. |
| Response time | Median time from sending the call to receiving the full answer |
| Before a run | The self-test of every task and scoring rule must pass. |
|---|---|
| Completeness | Every model has answered every task in all three runs. |
| Error rate | Below 5 percent per model |
| Outlier list | Answers where judges disagree by 2 points or more, models that moved more than 15 points since last week, and 10 percent randomly chosen answers |
| Approval | A person at Aprex reviews the outlier list and approves the run in the changelog. |
| The website | The site does not build if a run has invalid numbers, a missing work type or too high an error rate. |
Every run is published as a CSV file with one row per model, and the barometer page carries structured data (schema.org Dataset). The numbers are free to use with Aprex credited and a link to the barometer.
| Column | Content |
|---|---|
week | Week of the run, e.g. 2026-U41 |
model_id | Model ID in AI Gateway |
model | Model name |
provider | Provider |
tier | topp (top) or rimelig (budget) |
total | Total score 0–100 |
ci_low | Lower bound, 95% confidence interval |
ci_high | Upper bound, 95% confidence interval |
dokumentuttrekk | Work type score 0–100 |
kundesvar | Work type score 0–100 |
oppsummering | Work type score 0–100 |
regelverk | Work type score 0–100 |
klassifisering | Work type score 0–100 |
nynorsk | Work type score 0–100 |
cost_per_100_usd | Measured cost for 100 tasks, USD excl. VAT |
median_latency_ms | Median response time in milliseconds |
error_rate | Share of calls that failed after one retry, 0–1 |
No provider pays for or influences the barometer, and none sees results before they are published. Aprex itself uses models from several of the providers in the solutions we build.
Write to us if you want to know more, have ideas for tasks or have found something that should be fixed.
Email Aprex →