AI Models: Three Everyday Traps

A launch scheduled for 23:30 on 31 December in Kiribati’s UTC+14 time zone happens at 23:30 on 30 December in Hawaii, which is UTC−10. Seven of the eight tested models gave that answer. Claude Opus 4.8 instead wrote “2026-12-31, 03:30” for Hawaii—an error of 20 hours and a calendar day—even while correctly saying that no location had reached New Year at the moment of launch.
Disclosure: I work with OrcaRouter and used one key to access the models in this evaluation. OrcaRouter routes requests by task complexity, cost-effectiveness, and latency.
AI-generated illustration featuring official model logos; logos and model names are used descriptively and remain the property of their respective owners.
Compare the models in this article through OrcaRouter’s model catalog.
That is the point of these deliberately small puzzles. The arithmetic is not difficult in isolation. The trap is retaining the rule: a 24-hour offset means a whole-day change, not merely a time adjustment. The same dynamic applies when a percentage has a technical definition, or when a hypothetical legal convention changes which calendar date controls.
Three prompts, eight models, one selected response each
This was a prompt-specific case study of 24 selected gateway/API calls: three custom questions posed to eight models—Claude Fable 5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, and GLM-5.2. There was one selected answer per model-question cell, not a repeated-sampling study of stable behavior. All selected calls were recorded as successful, and web search was neither requested nor effective in these calls.
The questions tested stated conventions and arithmetic definitions. They were not requests for real legal or culinary advice. In particular, the leap-day scenario supplied its own rules for two fictional jurisdictions; the exercise was whether models followed them.
| Everyday trap | Observed result across the selected answers |
| Time-zone conversion | Seven answers gave Hawaii as 30 December 2026, 23:30; Claude Opus 4.8 gave 31 December 2026, 03:30. |
| Sourdough hydration | All eight calculated 416.7 g flour and 283.3 g water for a 700.0 g dough at 68% hydration. |
| Leap-day contract convention | All eight said 2018 was not a leap year, then applied 1 March in Jurisdiction A and 28 February in Jurisdiction B. |
The time-zone question: a nearly correct answer can still be wrong
The launch prompt fixed the event at 2026-12-31 23:30 in UTC+14. Converting that instant to UTC gives 09:30 on 31 December. Samoa at UTC+13 is therefore 22:30 on 31 December, while Hawaii at UTC−10 is 23:30 on 30 December. None of the three locations has yet reached midnight on 1 January; Kiribati gets there 30 minutes after launch. GPT-5.6 Sol laid out that chain explicitly.
GPT-5.6 Terra, GPT-5.6 Luna, Claude Fable 5, Grok 4.5, Gemini 3.5 Flash and GLM-5.2 reached the same local dates and times.
Claude Opus 4.8 recognized that Hawaii was 24 hours behind Kiribati, but its written subtraction produced “2026-12-31, 03:30,” not the prior day at 23:30. Its later explanation also treated Hawaii as 03:30 on 31 December. This is a useful reader-level warning: an answer can have the right headline—“none” celebrates New Year first—while containing a wrong intermediate fact that matters for scheduling.
The hydration question: following the definition, not just the percentage
The sourdough prompt began with 400 g of flour at 75% hydration. Every model first treated hydration as water relative to flour, yielding 300 g of water and a total dough weight of 700 g. Holding that total constant at 68% hydration means solving flour plus 0.68 times flour equals 700. The resulting amounts are 416.7 g flour and 283.3 g water.
This is perhaps the cleanest success in the set. The key wording was “TOTAL dough weight identical.” A weaker response could have simply reduced water while leaving flour at 400 g, preserving neither the stated total nor the requested new ratio. None of the selected answers made that mistake.

Reproducible data figure from this article’s selected API records; it is not a general reasoning ranking.
The leap-day contract question: the stipulated rule is the answer
For the final prompt, a person born on 29 February 2000 reaches 18 in 2018. The prompt specified that Jurisdiction A deems the birthday to fall on 1 March in non-leap years, while Jurisdiction B uses 28 February. Every selected response first noted that 2018 is not a leap year, then concluded that the contract matures on 1 March 2018 in A and 28 February 2018 in B.
The outcome is less about legal knowledge than instruction discipline. The model did not need to determine what any real jurisdiction does; it needed to apply the supplied convention consistently. That distinction matters because a polished answer can sound legally authoritative even when it is only operating within a hypothetical.
What this case study suggests—and what it does not
The observed pattern is narrow: all eight selected answers handled the percentage and leap-day definitions, while one selected time-zone answer made a substantial date-and-time conversion error. It would be a mistake to turn that into a general reasoning leaderboard. A single call is an observation, not proof of a model’s typical performance, and vendor effort labels should not be read as equivalent compute budgets across providers.
A reference-guided DeepSeek v3 judging process was used in the underlying evaluation, but it is not human professional review. The concrete answer records are more informative here than any aggregate score: readers can inspect exactly why the Hawaii output is wrong and why the other two exercises were accepted.
For broader context, Humanity’s Last Exam is a closed-ended academic benchmark, but it does not measure these custom everyday-rule prompts and should not be used to validate or rank this case study.
Practical takeaways
- Check the governing definition first. Is hydration water divided by flour? Does a contract scenario state a specific convention? The definition determines the calculation.
- Ask for intermediate values on date and time tasks. A UTC conversion and local-date table make an error easier to catch.
- Do not accept a correct final conclusion as proof that every detail is correct. The Claude Opus 4.8 response got the New Year conclusion right while misstating Hawaii’s local launch time.
- Treat these tools as drafts for consequential work. The examples were simplified convention tests, not legal or culinary advice.

Editorial illustration; it frames careful use of stated rules and is not test evidence.
Limitations
This covers three custom prompts with one selected answer per model-question cell. It tests stated conventions and arithmetic definitions rather than real legal or culinary advice. The results describe gateway/API behavior in the recorded calls, not consumer subscription products or general model quality. Finally, the reference-guided DeepSeek v3 judgments are not human professional review.
Explore the Models
Explore the current catalog on OrcaRouter Models.
This evaluation was run through OrcaRouter. The author works with OrcaRouter; model access does not imply affiliation with, endorsement by, or sponsorship from model providers.
Model names and logos are used descriptively. All trademarks belong to their respective owners.
Sources
- Humanity’s Last Exam — Center for AI Safety and Scale AI collaborators; retrieved 2026-07-17.



