ENUMA / EVALUATION SUITE
ENUMA-Bench
How well does a world follow your intent?
SIX INTERACTION CATEGORIES
Evaluating geometric, semantic, and appearance-level control in interactive world models.
01 / QUANTITATIVE EVALUATION
Results forthcomingLeaderboard
Unranked preview. Methods follow the draft report; scores have not been released.
— Not yet reported02 / QUALITATIVE EVALUATION
See the interaction.
Explore the task. Inspect the outcome.
GALLERY LAYOUT PREVIEW
Illustrative ENUMA demos are shown below. Official benchmark cases and matched model comparisons are forthcoming.
Examples are on the way.
Benchmark examples for this category will appear here.
03 / WHAT WE EVALUATE
A broader test of control.
Six categories, grounded in the ENUMA technical report.
Evaluation protocol
Scores use a 0–100 scale. Dataset counts, the final scoring procedure, and release details will accompany the benchmark.