← GSLF / ProjectsUNIMATRIx / 03

Models in a world of consequences

UNIMATRIx

Simulation as a Benchmark.

Evaluate language models through decisions in a simulated world. Fixed, versioned recipes set the conditions; the candidate model changes. Compare the resulting performance across eight domains.

Fixed recipesObservable decisionsComparable runs

↳ How it works

A recipe defines the cases, roles, seeds, peers and scoring rules. Models face the same scenarios, and their actions produce measurable outcomes across a sequence of simulation steps.

Completed runs retain their observations, responses and events for inspection. Rankings group results by recipe revision, engine fingerprint and compute reporting track, so comparisons share the same conditions.

↳ Explore the leaderboards

Choose the conditions for this comparison.

Loading recipe details…

Cases
Domains
Scripted peers

Loading illustrative leaderboards…

↳ Reading the scores

Scoring and comparison notes

Scores range from 0 to 100; higher is better. The aggregate gives each domain equal weight. Within a domain, levels are equally weighted, with metrics weighted according to the recipe. Every case must finish before a real run receives an aggregate score.

Standard V1 contains 192 cases; Compact V1 contains eight and has a separate leaderboard. The bundled recipes are experimental definitions. Compact V1 uses a single seed and cannot estimate variation across seeds.

This page uses fictional models and illustrative domain scores. Its overall demo score is the arithmetic mean of the eight domain scores. It does not represent a live benchmark run.

↳ Source

github.com/gslf/UNIMATRIxSource, benchmark recipes and run explorer.