Microsoft Test Finds Most Failed AI Agent Runs Show No Error
Get News Desk by email
One email a week: the AI stories that mattered, explained simply.
Anthropic's Claude Opus 5.5 scored highest, completing 67.16% of 507 office-style tasks on its first attempt.
Also today: GitHub removes four models from Copilot, and Germany's Aleph Alpha releases a free bilingual model.
Microsoft and Hugging Face released ThinkingBox on October 3, 2026. The test checks what an AI agent actually changed in a company's records, instead of trusting what the agent said it did. In their runs, about two-thirds of failed attempts ended cleanly, with no error message at all. Most of those (77.61%) had written a wrong value into a field. So an agent that says "done" may not be done. If you let an assistant update a spreadsheet or a customer database, check the result yourself. The researchers say about 80% of failures came from how agents handled their tools. Faulty reasoning caused far fewer. read the ThinkingBox post ↗
GitHub Retires Four AI Models From Copilot
GitHub switched off four models in Copilot, its coding assistant, on October 2, 2026: Claude Opus 4.7, Kimi K2.7 Code, Gemini 3.5 Flash and Gemini 3.6 Flash. They are gone everywhere in Copilot, from chat to autocomplete. GitHub's changelog names the replacements as Claude Opus 5.5, Kimi K3 and Gemini 3.8 Flash. Using Copilot at work and can't see the new ones? Your company's admin may need to turn them on first. GitHub changelog ↗
Aleph Alpha Releases Kolibri, a Free German-English AI Model
Aleph Alpha, Germany's best-known AI lab, published Kolibri-1 on Hugging Face on October 3, 2026, under the Apache 2.0 license, which allows commercial use. The model works in German and English, and 21.3% of its training text was German. That matters for European businesses that want an AI they can run on their own servers. It won't run on a laptop. The download is about 78GB. Aleph Alpha's own tests put it ahead of similar-sized models on math and science questions; no outside lab has checked those numbers yet. Aleph Alpha's announcement ↗
Update: OpenAI's Codex Lead Teases a Faster GPT-6.1 Sol
Thibault Sottiaux, who leads Codex at OpenAI, posted two short teasers on X early on October 4, 2026. OpenAI has not announced a release, a date or a price. Sol is the lower-cost model News Desk covered on October 1, which scored within a point of OpenAI's top model on one test at a fraction of the cost. A faster version would matter most to people who wait on Codex, OpenAI's coding assistant, to finish its work. see the post ↗
Meta Publishes Six Math Papers Written With Muse Spark
Meta said on October 2, 2026 that mathematicians used Muse Spark, its AI model, in the ordinary chat app to work on six open math problems. It published the resulting papers. According to Meta, people directed the work and a second group reviewed it. Each paper marks which passages the model drafted. No new model came with the announcement. Meta's research post ↗
Trending on GitHub
Agent-Reach Lets AI Assistants Read Reddit and YouTube
Agent-Reach gained 3,952 stars on GitHub this week. The free command-line tool lets an AI assistant search and read X, Reddit, YouTube and Chinese sites such as Bilibili, without paying for official data access. That gives a coding assistant a way to research what people are saying online. It reads the sites' pages directly, so a change on one of those sites can stop it working. Agent-Reach on GitHub ↗
ExplainerAgent-Reach Lets AI Agents Read Twitter and Reddit Without API Fees →StepFun, a Chinese AI lab, says the downloadable version of its Step 5 Preview model arrives October 15.
Sources
Get News Desk by email
One email a week: the AI stories that mattered, explained simply.
Build something with AI
Get the 16 prompts we use to plan, build and launch real projects with AI. Free.
Get the free prompts →