
Tech • AI • Robotics • Game
OpenAI, Mistral, Google, and Anthropic rolled out major AI updates, led by OpenAI’s release of 722 math manuscripts generated from an unreleased frontier model tested on about 4,000 research problems.
OpenAI released 722 AI-generated mathematical manuscripts spanning 372 result families after evaluating an internal model on roughly 4,000 research problems. The work reaches beyond standard benchmarks into areas including the Riemann zeta function, Hodge conjecture-related questions, spin glasses, quantum magnetism, and operator algebras. Some outputs include Lean-formalized proofs, allowing parts of the work to be checked by proof software rather than accepted on trust alone.
The release does not amount to proof that AI has solved landmark open problems such as the Riemann Hypothesis or the full Hodge Conjecture. However, the scale is notable: OpenAI considered the findings substantial enough to consult the Institute for Advanced Study, independent mathematicians, and AI advisory groups on how to share them. OpenAI also indicated that successful results required only about three hours of ChatGPT Pro-equivalent compute on average, suggesting a sharp rise in the efficiency of machine-assisted mathematical exploration.
On the second day of a 28-day shipping campaign, OpenAI introduced four user-facing updates. Auto-review is now free for signed-in ChatGPT users and no longer counts against normal usage, after previously consuming an estimated 2% to 10% of plan capacity. The company also simplified API usage tiers into Build, Launch, and Grow, with the top tier now reached at $500 in cumulative API payments rather than $1,000.
OpenAI added a meetings feature for the ChatGPT macOS desktop app in beta for Pro and Business users. It can capture meeting notes, create summaries and next steps, and store them inside workspace context for follow-up actions. The new Decision API gives developers a fast classifier for choosing which model, tool, or action to use next, with OpenAI saying decisions can run up to 10 times faster than routing through a larger model in the Responses API.
Mistral launched Mistral Large 4, also described as Le Chat, a natively multimodal one trillion-parameter model using a mixture-of-experts design with 49 billion active parameters at a time. The company positioned it as a leading open-weight model developed in the US or Europe, with a particular focus on enterprise and regulated sectors such as cybersecurity, finance, and manufacturing. API access is available now, while open weights are expected later this month.
Mistral said the model solved 18 of 19 cybersecurity capture-the-flag challenges and performed strongly in malware reverse engineering. In one test, it identified an unknown sample as Cobalt Strike, extracted its configuration and indicators of compromise, and produced a full investigation report with a YARA rule in about 12 minutes, a task that can take a human analyst much of a working day. The model also offers a 1 million-token context window, though speed and cost tradeoffs may limit its appeal as an everyday coding assistant.
Google launched Nanabanana 2.1, improving image generation in visual design, mask-based editing, subject consistency, and rendering speed. In one cited comparison, it generated images at about 4 cents each in 11 seconds, versus roughly 7 cents and 50 seconds for a competing model. Google also released Embedding Gemma 2, a 740 million-parameter open multimodal embedding model for on-device use across text, images, audio, video, and code under the Apache 2.0 license.
Anthropic opened applications for Mythos 5.1, a restricted model reportedly sharing weights with another 5.1 system but operating with looser cybersecurity and biological safety limits for approved users. Access is structured in tiers including defense, red team, and a more restricted specialized level for organizations that clear deeper reviews. The company also expanded integration between Claude and Google Docs, Sheets, and Slides, and added finer effort controls for Claude Code sub-agents.
A developer reported using Opus 5.5 to recreate seven major Adobe-style applications as free, open-source tools written in Rust, including a Photoshop alternative. The projects are described as early and not yet full professional replacements, but they illustrate how advanced coding models may lower the cost and team size required to build serious software products. The trend adds to evidence that frontier models are beginning to affect both research workflows and commercial software development.
The latest releases show AI moving simultaneously deeper into scientific research, enterprise automation, cybersecurity, and software creation. The central question is no longer whether these systems can assist experts, but how quickly their output can be verified, governed, and turned into reliable real-world work.
Ask a question