Tech • AI • Robotics • Game

VIDEO
ENFR

Opus 5.5 turned my company into a video game

6/10
AIRenaud DékodeOctober 10, 2026 at 02:00 PM58:35
Audio player
0:00 / 0:00

TL;DR

A hands-on comparison of Claude Opus 5 effort settings found that medium and extra delivered the best balance of quality, speed and cost when turning a structured work archive into a playable management game.

KEY POINTS

A game built from a work archive

The test asked Claude Opus 5 to create a small playable game from a structured professional file system described as a vault. The game reproduced a real work environment, including colleagues, recurring tasks, content planning, invoices, interviews and studio operations. The model inferred details such as names, roles, branding colors, schedules and workspace layout from existing files rather than from a highly detailed prompt.

The goal was to test reasoning, creativity and cost

The exercise was designed as a one-shot benchmark across all available effort modes in Claude Opus 5, with the same prompt used each time. The comparison focused on output quality, reasoning depth, verification steps, generation time, token usage and estimated dollar cost. It also tested whether the model could search through prior context and organize scattered information into a coherent product.

Low effort produced a workable but limited result

In low mode, the model generated a functional isometric game in 12 minutes 39 seconds. It used about 78,000 standard output tokens and roughly 5 million cached tokens, for an estimated cost of $3.82. The result included missions, characters, a basic interface and recognizable business details, but movement and interaction were simpler, with tile-based navigation and more visible rough edges.

Medium effort showed a major quality jump

In medium mode, the game became significantly more polished, with freer movement, stronger visual presentation, better interface cues and more immersive touches such as music linked to specific characters. The model also recreated studio elements and environmental details with greater fidelity. That version took nearly 30 minutes, cost $9.94, used about 186,000 standard tokens, and expanded to 57 reasoning steps, making the leap from low to medium the most striking improvement in the test.

Extra appeared to offer the best value for serious projects

The comparison concluded that extra provided the strongest quality-to-price ratio for projects that needed to be more than a quick prototype. It was described as the setting most likely to be used for production work after initial experiments in medium mode. The output remained fluid and usable while showing a stronger grasp of structure, polish and execution.

Max favored heavier coding over usability

The max setting pushed further into raw development complexity, adding more advanced behavior and broader technical ambition. But that came with trade-offs: longer runtimes, around 50 percent more token use than lower advanced modes, and a result that felt less elegant to use. For web apps or lighter interactive tools, the extra complexity looked unnecessary, even if the underlying code generation was more ambitious.

Ultra is not above max

The test also highlighted confusion around ultra mode. It was described not as a level above max, but as something closer to extra with mandatory sub-agent orchestration and parallel task handling. In this benchmark, ultra landed very close to extra in time, cost and quality because the assignment was a single end-to-end build rather than a workflow that benefited from parallel subtasks.

Cached tokens reduced the real cost

One notable takeaway was that large token counts can overstate actual expense. Much of the work was billed as cached tokens, which are cheaper because the system reuses repeated context and intermediate material during iterative generation and self-checking. That is why a project that appeared to consume millions of tokens still remained under $10 in some modes.

Structured files mattered as much as prompting

The quality of the result depended heavily on having a well-organized archive of folders and documents. The test argued that users who keep a clear working repository of clients, invoices, ideas, formats and strategy notes can get much richer outputs from AI tools. In that setup, the model did not just generate code; it turned the archive itself into a source of world-building, mechanics and task design.

Deployment tools were presented as the next step

For users wanting to publish AI-generated projects, the workflow pointed to Hostinger Connector, which links hosting controls to coding agents through an MCP integration. The idea is to let an AI assistant deploy, update and manage projects directly on a domain or subdomain without requiring advanced technical knowledge. Entry-level hosting was presented at €3.59, with the integration positioned as a way to remove friction between prototype and launch.

CONCLUSION

The benchmark suggests that Claude Opus 5 is most useful when paired with a rich, organized personal knowledge base and used to augment human creativity rather than replace it. For most users, medium and especially extra appear to be the most practical settings for turning real work context into functional software.

Ask a question

More from AI