
Tech • AI • Robotics • Game
OpenAI published 722 AI-generated mathematics manuscripts on October 6, grouped into 372 result families after testing an internal frontier model on roughly 4,000 research problems. The release spans topics including the Riemann zeta function, Hodge-related questions, spin glasses, quantum magnetism, and operator algebras, with 235 families also accompanied by Lean formalizations. The company says the papers are provisional and may contain errors, but some claims were considered significant enough to involve the Institute for Advanced Study, outside mathematicians, and AI advisors before publication. A key datapoint is efficiency: successful results reportedly used about three hours of ChatGPT Pro-level compute on average, suggesting a sharp jump in machine-assisted mathematical search.
Mistral unveiled Mistral Large 4, its first new flagship in nearly a year, built as a mixture-of-experts model with 1 trillion parameters and about 49 billion active per query. The model takes text and images as input, outputs text, supports reasoning and code generation, and offers a 1 million-token context window aimed at enterprise workloads. Mistral is pitching sovereignty as a core differentiator, saying the system is hosted in Europe and trained entirely in France on its own infrastructure. That makes the launch as much a geopolitical and compliance play as a benchmark contest, especially for regulated sectors in finance, industry, and cybersecurity.
Anthropic launched Claude Haiku 5.5 as its fastest and cheapest small model, cutting base pricing to $0.10 per million input tokens and $0.50 per million output tokens for prompts under 100,000 tokens. That is roughly a 90% reduction from Haiku 4.5, with the company claiming average operating cost about 75% lower overall. Anthropic also reported large benchmark gains, including jumps on OSWorld, Humanity’s Last Exam, and Chartography, positioning the model for coding sub-agents, automation, summarization, and customer support. The move sharpens price pressure on high-volume enterprise inference, where cheaper small models increasingly handle routine work once reserved for larger systems.
A leaked Anthropic IPO filing outlined explosive top-line growth alongside extraordinary paper losses, making it one of the clearest financial snapshots yet of a frontier AI lab. Reported revenue rose from $386 million in 2024 to about $4.6 billion in 2025, with annualized recurring revenue near $65 billion by late summer and a projected $100 billion run rate by year-end. The eye-catching $42 billion loss was said to be driven largely by accounting treatment of convertibles granted to Google and Amazon, with roughly $34 billion reflecting dilution rather than direct cash burn. Wall Street is likely to focus less on the headline loss than on whether such growth can outrun infrastructure spending and long-term model costs.
OpenAI introduced Spaces and Pages inside ChatGPT, adding a shared workspace layer for documents, images, visualizations, recurring briefs, and collaborative editing. Spaces replaces the old library by gathering assets from chats and connected sources such as Google Drive, while Pages provides a live document editor for long-form work. Users can request line edits with Ask for change, generate inserted sections, and use the main chat to restructure an entire document without leaving the app. The rollout, initially aimed at Pro and Business users, signals ChatGPT’s push from single-thread assistant to multi-document team workspace.
OpenAI broadened Codex with Codex Cloud, plus integrations for Slack and Microsoft Teams, to support long-running coding agents across desktop, web, terminal, and collaboration tools. The cloud service lets developers assign tasks to managed environments instead of keeping local machines or improvised remote boxes alive for hours or days. Codex can inspect repositories, resolve dependencies, prepare startup scripts, and generate reusable operating guides, reducing the setup burden for persistent agent work. Related demonstrations showed the product correlating Grafana telemetry, deployment context, and code changes to propose incident-response patches within minutes while leaving human approval in the loop.
Grockbot is being repositioned as a routing layer that can dispatch tasks to outside systems such as Claude Opus 5.5, MidJourney, and Suno instead of relying only on Grok models. The shift follows criticism that Grok trails top rivals on hard coding and reasoning workloads, making best-of-breed orchestration a more practical strategy than single-model loyalty. In effect, users could stay inside one agent interface while different providers quietly handle code, image generation, audio, or research behind the scenes. If broadly adopted, that model could make model choice increasingly invisible and move competition toward orchestration, interface control, and agent reliability.
A parallel theme across launches was the rise of governance layers for persistent software agents, from OpenClaw Enterprise to Qodo and enterprise deployments at Stack Overflow. These systems focus less on raw generation and more on multi-tenant controls, security boundaries, access rules, organizational memory, and review workflows that let teams share and supervise agent work. The enterprise problem is no longer just producing code faster; it is proving who can see what, which changes are safe, and how accumulated engineering knowledge is carried into automated decisions. That emphasis suggests the next battleground is not only model quality, but operational trust, compliance, and control over autonomous workflows.