Tui

Evaluating frontier LLMs on operational restaurant management

Right now, frontier models are being benchmarked on metrics that are not tied to economic value. So as a result, agents running on these frontier models are lackluster in completing tasks that are non-deterministic inside of enterprises. We think it is because of two reasons:

1evals test for deterministic tasks while real work is not
2behavior data on how humans operate is valuable for alignment but scarce and needs to be collected from robust environments.

So we decided to construct a gamified restaurant environment where frontier models act as shift manager. Chef, supply management, host, and waiter are fixed worker policies that are simulating employees. The model only decides when to intervene on equipment failures, allergen verification, and other notable crises. Every model runs the same nine episodes. We report a deterministic outcome score and a separate trajectory review using an Archipelago-style harness grader. The same environment also lets us collect human behavior traces, so we can see whether models get a sense of business from the game and use that data for post-training on management, negotiation, and other capabilities that carry economic value but resist standard task formats.

Four colorful agents coordinating a restaurant shift around a table

Results

Sort by

Swipe for more columns

RankModelAll-model (0–1)Composite ScoreRubricPass@1ServiceResourcesRecoverySafetyTokens / EpAction
1
Muse Spark 1.3
0.786 / 1.000
85.0 / 100
39/54 (0.722)66.7%89.8%98.4%50.0%100%10,352
2
DeepSeek V4 Pro
0.769 / 1.000
85.2 / 100
37/54 (0.685)66.7%89.8%99.1%50.0%100%10,615
3
Grok 4.6
0.761 / 1.000
85.5 / 100
36/54 (0.667)66.7%89.8%100.0%50.0%100%11,225
4
GPT-5.6 Terra
0.741 / 1.000
85.2 / 100
34/54 (0.630)66.7%89.8%99.1%50.0%100%4,962
5
Gemini 3.8 Flash
0.740 / 1.000
85.0 / 100
34/54 (0.630)66.7%89.8%98.4%50.0%100%7,123
6
Claude Opus 5
0.732 / 1.000
85.2 / 100
33/54 (0.611)66.7%89.8%99.1%50.0%100%8,202
7
Nemotron 3.5 Lightning
0.725 / 1.000
78.3 / 100
36/54 (0.667)55.6%85.9%93.5%33.3%100%14,664
8
Claude Sonnet 5
0.721 / 1.000
85.0 / 100
32/54 (0.593)66.7%89.8%98.4%50.0%100%8,151
9
GPT-5.6 Sol
0.721 / 1.000
85.0 / 100
32/54 (0.593)66.7%89.8%98.4%50.0%100%4,970
10
Kimi K3
0.713 / 1.000
85.2 / 100
31/54 (0.574)66.7%89.8%99.1%50.0%100%6,013
11
Fable 5.1
0.712 / 1.000
85.0 / 100
31/54 (0.574)66.7%89.8%98.4%50.0%100%8,461
12
GPT-5.6 Luna
0.631 / 1.000
78.0 / 100
26/54 (0.481)55.6%84.8%93.5%33.3%100%5,172
13
Gemma 4 31B
0.630 / 1.000
74.1 / 100
28/54 (0.519)0.0%89.8%99.1%50.0%89%8,126
14
DeepSeek V4 Flash
0.561 / 1.000
79.0 / 100
18/54 (0.333)55.6%86.5%94.9%33.3%100%13,539
15
Gemini 3.7 Flash
0.502 / 1.000
56.0 / 100
24/54 (0.444)44.4%89.8%98.4%50.0%67%5,002

Key findings

We ran thirteen frontier models through the same nine episodes, then regraded every trajectory under the six-criterion rubric, which is 54 qualitative judgments per model. A few things came out of the runs that we did not expect going in, and widening the rubric from two criteria to six changed what we could see.

1

models that get the same result do not manage the same way

Eight of thirteen models land between 85.0 and 85.5 composite, and all eight post the exact same 89.8% service rate. Outcomes alone would call that group equivalent. The six-criterion trajectory review pulls the same group apart from 31 of 54 criteria up to 39. Muse Spark 1.3 and Fable 5.1 run the shift to the same deterministic score, but one passes 39 of 54 rubric criteria and the other passes 31, so one of them is making the right calls and the other is arriving at a similar number by accident. Widening the rubric from two criteria to six is what exposes that spread: the gap shows up in coordination, sequencing, and monitoring, not in the outcome numbers. This is exactly what the trajectory review is there to catch, and it is the part that carries over to real operational work.

2

nobody handles saturday night well

The 'Saturday Night' scenario is where we found every single model to struggle the most. Every model we ran scored 0 on recovery there, and composite fell to around 58 across the board. The old two-criterion rubric could tell us that models failed there but not how, and the new criteria name it: across the 30 Saturday Night episodes with per-episode verdicts, Coordination and sequencing passes 0 times, and Monitoring and adaptation passes twice. The same two criteria pass 30 of 30 and 26 of 30 on Quiet Tuesday. So it is not that the models stop caring about the shift under load, it is that they stop tracking it. They respond to whichever problem just appeared instead of keeping the rest of the floor moving, which is what we suspected before and can now point at.

3

safety is where models in the middle of the pack perform worse on

Service is nearly identical across models, so the thing that decides a 56 from an 85 ends up being a single missed allergen check. Gemini 3.7 Flash posts the same 89.8% service rate as the top of the leaderboard and still finishes last. Safety is a gate and not a metric, so a model cannot serve its way out of it, which is how we think it should work in an operational setting.

4

spending more tokens does not buy better management

There is no relationship in our runs between how much a model thinks and how well it runs the restaurant. Nemotron 3.5 Lightning spends the most tokens of any model in the field and finishes 6th of 13, and DeepSeek V4 Flash is second in spend and finishes 12th. GPT-5.6 Terra spends the fewest tokens, about a third as many as Nemotron, and finishes 4th. So whatever operational judgment is here, it is not something the models are buying with more tokens.

5

no model has passed all three shifts

Pass@1 tops out at 66.7%, which is two of the three shifts and never the third. This is not noise or provider failures, because every published model completed all nine episodes. It is the same shift failing the same way for everyone, and that is exactly the kind of task we want human behavior traces on.

6

models are bad at the big picture

Sorting the six criteria by how often they pass gives a clean ordering. Across the 90 episode reviews from the ten models graded end to end, Constraint discipline passes 96.7% of the time and Recovery handling 84.4%, so models reliably respect the allergen rules and reliably fix a failure once it is in front of them. Everything that requires holding the shift in mind is much worse: Resource proportionality passes 52.2%, Coordination and sequencing and Monitoring and adaptation 43.3% each, and Decision quality 40.0%. The split widens with load. Constraint discipline barely moves, holding 96.7% even on Saturday Night, while Coordination and sequencing falls from 100% on Quiet Tuesday to 30% on Dinner Rush to 0% on Saturday Night. Models are not trading safety away when the floor gets busy. They are losing track of the floor.

Methodology

This benchmark evaluates whether a model can make operational decisions under changing demand, cost, and safety constraints in a controlled environment. Outcomes and decision traces remain inspectable so researchers can study both performance and failure modes.

Constants for the environment

Each episode runs for 120 simulation ticks. The supporting crew follows fixed policies at the same stations:

The entrance handles the queue and allergen screening; dining handles expedite and handover; the kitchen handles grill, fryer, and assembly.

All models see the same arrivals, tickets, and interventions. Options are not labeled as safe or risky, keeping the comparison focused solely on management decisions rather than variation in worker behavior.

What each environment tests

Quiet Tuesday

Low arrival volume, no scheduled equipment failure, no supply shocks. With almost nothing forcing a Recovery event, the composite is carried by Service and Resources, so this case isolates whether a model overspends or overstaffs for pressure that never shows up. The allergen screen still runs on every eligible order regardless of how quiet the floor is, so a missed check here reads as a genuine Safety gate failure, not a rushed decision under load. The harness's Decision quality and Resource proportionality criteria are watching for the model treating a calm night as if it were busy, and Resource proportionality is the criterion models fail most often here, passing 46.7% of the time while every other criterion except Decision quality passes at least 86%.

Dinner Rush

Higher arrival volume with one scheduled equipment failure partway through the shift. This is the first case where Recovery is actually scored. The model has to notice the failure, reroute the affected orders, and bring the backlog back down rather than letting it compound for the rest of the shift. All six harness criteria apply here. Decision quality is scored on the reroute itself, Recovery handling on whether the disruption was actually resolved rather than just acknowledged, and Coordination and sequencing and Monitoring and adaptation on whether the rest of the shift kept moving while the failure was being handled.

Saturday Night

Peak arrival volume with multiple supply shortages and station bottlenecks landing close together instead of one at a time. Service, Resources, and Recovery are all under pressure simultaneously, so this case tests whether a model can triage competing failures instead of fully resolving whichever one appeared most recently. It is also the case where every model we ran scored 0 on Recovery, and where Coordination and sequencing passes zero times out of thirty. See findings 2 and 6 below.

Evaluation and scoring

The benchmark reports two complementary signals: a deterministic outcome score and a trace-based trajectory review. The first measures what happened; the second examines whether the model's decisions fit the situation.

Deterministic outcome score

The outcome score is computed directly from the saved event trace. It does not use an LLM, network calls, randomness, or the current clock, so replaying the same trace produces the same score.

Replay validation confirms that each episode:

  • the run uses the frozen Fixed Workers v2 contract;
  • the case ID, 120-tick duration, guest arrivals, and scheduled equipment failures match;
  • the cash ledger is internally consistent;
  • orders are created and served in a valid sequence;
  • the episode reaches a valid terminal state;
  • replaying the trace reproduces the recorded grade exactly.

The validated trace is then used to calculate:

  • Service: percentage of eligible orders served by their deadlines;
  • Resources: normalized net operating return relative to the published floor and ceiling;
  • Recovery: how quickly a disruption is repaired and the affected order is recovered, when the case includes a recovery event;
  • Safety: 100 when there are no safety violations, otherwise 0 since this criterion has life-or-death implications
  • Composite: the mean of the applicable operational metrics, multiplied by the Safety gate.

Safety is intentionally a hard gate. A model cannot compensate for a safety violation with faster service or better revenue.

Archipelago-style harness grader

The second layer runs a structured harness over the model's saved trajectory, separately from the deterministic outcome score. It runs only on valid, replay-verified episodes and evaluates decisions in the context of the restaurant state.

The first version of this rubric scored two criteria. That was enough to show that models reaching identical outcomes were not managing identically, but not enough to say how they differed, and it left too many models tied. The current restaurant-fixed-workers-archipelago/v2 rubric scores six, which regraded all 135 trajectories and separates rule-following from floor management. Deterministic scores were not touched by the regrade.

For each included episode, the harness uses:

  • the model's recorded actions;
  • relevant player actions and corrections;
  • equipment failure and restoration events;
  • the terminal observation;
  • the deterministic grade as reference evidence.

The harness scores six criteria under the restaurant-fixed-workers-archipelago/v2 rubric. The first two carried over from v1; the other four were added to separate models that reach the same outcome by different means:

  1. Decision quality — whether the selected actions fit the restaurant state and available information.
  2. Recovery handling — whether disruptions were identified and addressed appropriately.
  3. Constraint discipline — whether explicit safety and task constraints, including allergy requirements, were respected without unsupported overreaction.
  4. Resource proportionality — whether purchases, wasteful actions, and other interventions were proportionate to observed demand, capacity, and cost.
  5. Coordination and sequencing — whether actions across roles were coordinated and sequenced to keep work moving without avoidable conflicts.
  6. Monitoring and adaptation — whether recorded state changes were noticed and later actions adapted, including incidents, arrivals, and equipment conditions.

Each criterion receives a pass/fail verdict, rationale, and evidence ticks, so nine episodes produce 54 qualitative judgments per model. The review cites concrete trace ticks and does not recompute or reinterpret the deterministic score. Every verdict is readable per episode in the Rubric tab of each model's trace inspector.

Published all-model score

The published all-model score combines the two signals across nine valid episodes:

deterministic score = mean episode Composite / 100
trajectory score = passed rubric criteria / total rubric criteria
all-model score = (deterministic score + trajectory score) / 2

A model appears on the ranked leaderboard only when it has nine valid completed episodes, 54 reviewed trajectory criteria, and replay evidence for every included episode. Provider failures, malformed responses, and incomplete runs are recorded separately rather than treated as model behavior.

Pass@1 (All Requirements Met): true only when the shift reaches a valid terminal state, Safety is 100%, each operational score is ≥ 80%, and the final performance report is accurate.