Norwegian AI Barometer / Method

How the barometer is built.

This page is for anyone who wants to understand or check the numbers. It covers why we test, the principles behind the test, exactly how answers are scored and what the numbers cannot be used for.

Background

Why a Norwegian barometer?

The well-known language model tests mostly measure English, coding, maths and knowledge questions. They say little about what Norwegian businesses actually use AI for: reading an invoice, answering a customer, summarising a meeting or writing in Nynorsk.

When Aprex builds AI solutions, we have to choose a model for each task. That takes numbers on Norwegian work, not English exam questions. The barometer is the same measurements done systematically and published, so others can use and check them.

Models change from month to month, and new ones keep arriving. So the test runs every week with the same tasks, and you can follow how the models develop over time.

Foundation

Eight principles behind the test.

  1. 01

    Work, not trivia

    Tasks resemble what employees do with documents, emails and minutes, and each has an answer that can be checked.

  2. 02

    Key before the run

    Answer keys and scoring rules are written before any model answers, and they do not change after we see the results.

  3. 03

    Equal conditions

    Same instructions, same settings and same service for every model.

  4. 04

    Hidden task set

    Most tasks are never published, so they cannot end up in training data and give some models a head start.

  5. 05

    Automatic first

    Scores come from fixed rules wherever possible. Judge models are used only where quality must be rated, and count for at most 40 percent.

  6. 06

    Uncertainty shown

    48 tasks show direction, not decimals. The confidence interval shows how certain the differences are.

  7. 07

    Cost kept apart

    Score, cost and response time sit side by side and are never mixed.

  8. 08

    A person decides

    No run is published before a person has reviewed the outliers.

Task set

48 tasks, six work types.

Each of the six types has eight tasks. Two per type become public with their answer key, so anyone can see what the models get. Six are published now, and the rest follow when the task set is complete. The other 36 stay hidden. The tasks are in Norwegian.

TypesDocument extraction, Customer replies, Summaries, Regulation, Classification, Nynorsk
ContentFictional companies, people and documents, or public sources. No customer data, not even anonymised.
LanguageThe Nynorsk type is written in Nynorsk. At least one task in each other type has Nynorsk content.
RegulationAnswer keys are checked against the statute on Lovdata before each new version of the task set.
Each task hasTask text, answer key, a correct answer in other words, a deliberately wrong answer and scoring rules.
CheckThe key and the reworded answer must score at least 95, and the wrong answer at most 40, before a task can be used.
Run conditions

The same conditions for everyone.

ServiceVercel AI Gateway, one shared access point to every provider. The exact model ID is logged.
Instructions«Du løser en arbeidsoppgave for en norsk bedrift. Følg instruksjonene i oppgaven nøyaktig, og svar bare med det oppgaven ber om.» (Norwegian: you solve a work task for a Norwegian business; follow the task exactly and answer only what it asks.)
Temperature0 where the model supports it
ReasoningProvider default. The level is logged.
Tools and web searchOff
Max answer length2000 tokens
Time limit120 seconds
Repeats3 per task
ErrorsOn network errors or overload we retry once. If it fails again, the answer scores 0 and counts as an error.
Hidden tasksSent only to providers that do not train models on the data.
Scoring

How an answer earns points.

Field check

Used for document extraction and classification. The answer must contain a JSON object. Each field is compared with the key after normalisation: dates become YYYY-MM-DD, amounts become numbers so «15 625,00» and «kr 15 625,-» count the same, KID and other digit fields are compared without spaces, and categories ignore case. The score is the share of correct fields. Without JSON, the score is 0.

Checklist

Used for customer replies, summaries, regulation and Nynorsk. The answer earns points for each weighted item that must be included, and each item can be accepted in several forms, such as «29. september», «29.09» and «29/9». Then points are deducted for:

  • forbidden claims, such as promising compensation that was not agreed: the deduction is set in each task
  • answers over the word limit: −20
  • words from the wrong written standard: −10 per word, at most −50

Judges

Also used for customer replies, summaries and Nynorsk. Two models from providers other than the one being tested rate the answer without knowing who wrote it. They give 0–5 points per criterion on a fixed rubric. Current judges are Claude Opus 5.5, GPT-6 Sol, Gemini 3.1 Pro. The task score then becomes:

score = checklist × (1 − 0.4) + (judge mean ÷ 5 × 100) × 0.4

The judge instruction is the same in every run (in Norwegian):

Du vurderer et svar på en norsk arbeidsoppgave. Du får oppgaven, et fasitsvar og svaret som skal vurderes. Du vet ikke hvilken modell som skrev svaret. Gi 0–5 poeng per kriterium: 0 = mangler helt, 3 = brukbart med tydelige svakheter, 5 = like godt som fasit eller bedre. Svar bare med JSON på formen {"kriterier": [{"id": "...", "poeng": 0, "begrunnelse": "én setning"}]}.

Before the first publication we rate twelve answers ourselves and compare with the judges. If they differ by more than one point on average, the rubric is adjusted before numbers are published.

From task to total

How the table numbers are calculated.

Task scoreMean of the three runs
Type scoreMean of the eight tasks in the type
Total scoreMean of the six types. All types count equally.
Confidence intervalWe resample the tasks with replacement 10,000 times and compute the total each time. The interval is the middle 95 percent. Models with overlapping intervals count as equally good.
StabilitySpread between the three runs
CostActual price per call, including hidden reasoning, scaled to 100 tasks. Given in US dollars excluding VAT, because providers bill in dollars.
Response timeMedian time from sending the call to receiving the full answer
Quality control

Nothing is published automatically.

Before a runThe self-test of every task and scoring rule must pass.
CompletenessEvery model has answered every task in all three runs.
Error rateBelow 5 percent per model
Outlier listAnswers where judges disagree by 2 points or more, models that moved more than 15 points since last week, and 10 percent randomly chosen answers
ApprovalA person at Aprex reviews the outlier list and approves the run in the changelog.
The websiteThe site does not build if a run has invalid numbers, a missing work type or too high an error rate.
Versions

Numbers are compared only within one task set.

The task set is versioned. We make at most two new versions a year, and in the week a new version starts, we run both the old and the new so the trend can be linked. Changes to tasks, scoring, judges and the model list go in the changelog on the barometer page.

Data

Download and use the numbers.

Every run is published as a CSV file with one row per model, and the barometer page carries structured data (schema.org Dataset). The numbers are free to use with Aprex credited and a link to the barometer.

ColumnContent
weekWeek of the run, e.g. 2026-U41
model_idModel ID in AI Gateway
modelModel name
providerProvider
tiertopp (top) or rimelig (budget)
totalTotal score 0–100
ci_lowLower bound, 95% confidence interval
ci_highUpper bound, 95% confidence interval
dokumentuttrekkWork type score 0–100
kundesvarWork type score 0–100
oppsummeringWork type score 0–100
regelverkWork type score 0–100
klassifiseringWork type score 0–100
nynorskWork type score 0–100
cost_per_100_usdMeasured cost for 100 tasks, USD excl. VAT
median_latency_msMedian response time in milliseconds
error_rateShare of calls that failed after one retry, 0–1
Limits

What the barometer cannot tell you.

  • Results apply to these 48 tasks. A model may be stronger or weaker on your own tasks.
  • Judge models can be wrong and may prefer a certain style. That is why they count for at most 40 percent, come from other providers and are checked against human ratings.
  • Running through an API with fixed settings is not the same as using ChatGPT, Claude, Gemini or Copilot in the browser, where the provider adds its own instructions and tools.
  • Models are updated often, sometimes without a new model ID. A result applies to the week it was measured.
  • 48 tasks are enough to see clear differences, but not small ones. Check the confidence interval before you conclude.
  • Temperature 0 makes answers stable but not always identical. That is why each task runs three times.
FAQ

Technical questions

Why temperature 0?
It gives the most stable answers, so differences come from the model and not chance. Some reasoning models do not allow changing temperature. Then the default is used and logged.
Why three runs per task?
Answers vary slightly even at temperature 0. Three runs give a more reliable mean and show how stable each model is.
Why not let an AI rate every answer?
AI judges are fast but can be wrong and prefer their own style. Fixed rules can be checked by anyone. So judges are used only where quality must be rated, and count for at most 40 percent.
Why is web search off?
We measure what the model does with the information it is given. With web search, results would depend on the search engine, the timing and which pages exist.
How are reasoning models handled?
They run at the provider's default reasoning level. Hidden reasoning is included in both cost and response time.
Can I repeat the test myself?
Yes, with the public tasks. Task text, answer keys and the full method are open, so you can run them against the same models and compare. The hidden tasks are not shared.
I found an error. What do I do?
Write to post@aprex.no. We fix errors in tasks or method and record the change in the changelog.
Independence

No provider influences the numbers

No provider pays for or influences the barometer, and none sees results before they are published. Aprex itself uses models from several of the providers in the solutions we build.

Questions about the method?

Write to us if you want to know more, have ideas for tasks or have found something that should be fixed.

Email Aprex →