AI Digest

Daily AI update Digest

Curated updates from GitHub Copilot, Claude Code, Cursor, OpenAI Codex, and more.

486 items

Claude Code v2.1.288

New capability
Sun, Oct 4, 2026Claude Code GitHub Releases
  • Adds configurable `/code-review --max-findings` limits and built-in `gh api` support for cloud sessions without GitHub CLI.
  • Fixes session resilience issues around timeouts, auto-compaction, resume state, long prompts, and MCP tool calls.
  • Fixes a dangerous `rm`-via-shell path that could bypass prompting under permissive modes.
  • Adds safer hook behavior: unparseable PreToolUse or PermissionRequest hook inputs now block the call.

The Agent Said It Was Done. The Database Disagreed.

New capability
Sun, Oct 4, 2026Hugging Face Blog
  • Microsoft ThinkingBox is now available through Hugging Face for evaluating agents against isolated MCP tool sessions.
  • It grades resulting backend state and side effects instead of trusting an agent’s completion message.
  • Repeated runs expose whether an agent completes operational tasks reliably rather than succeeding once.
  • Useful for teams building tool-using agents that modify records, systems, or external services.

A model guide for the GPT-6 family

New capability
Sat, Oct 3, 2026OpenAI Blog
  • Covers GPT-6 model selection, reasoning effort, prompting, skills, tool use, and production workflow design.
  • Gives teams guidance for matching task complexity to model capability and cost.
  • Focuses on building reliable agent workflows, not one-off prompt experiments.
  • Useful baseline for teams migrating coding-agent workloads to GPT-6.

Copilot code review: API support and new default effort level

New capability
Sat, Oct 3, 2026GitHub Changelog
  • Copilot code reviews can now be requested through REST and GraphQL APIs.
  • API requests can set review effort, enabling review triggers from internal tools and CI workflows.
  • Balanced is now the default effort level for new and existing repositories using default settings.
  • Lite remains available through personal, repository, organization, or enterprise settings.

Selected models in GitHub Copilot deprecated

Migration recommended
Sat, Oct 3, 2026GitHub Changelog
  • Copilot retired Gemini 3.5 Flash, Gemini 3.6 Flash, Kimi K2.7 Code, and Claude Opus 4.7.
  • Suggested replacements: Gemini 3.8 Flash, Kimi K3, and Claude Opus 5.5.
  • Change applies across Copilot Chat, inline edits, ask mode, agent mode, and completions.
  • Enterprise admins may need to enable replacement models through Copilot model policies.

Model Router Benchmarks

Nice-to-have
Sat, Oct 3, 2026OpenRouter Blog
  • OpenRouter launched a benchmark page comparing model routers on quality, speed, and cost.
  • Benchmarks include routers and direct models, helping teams assess routing tradeoffs against fixed-model deployments.
  • Routing can reduce agent cost by assigning simpler work to cheaper models.
  • Cache rebuilds, routing latency, and poor task-complexity signals remain key risks for multi-model agents.

Codex CLI 0.160.0

New capability
  • Adds opt-in Guardian review context from prior instructions and agent handoffs.
  • Enables policy-permitted sessions outside a project using workspace defaults and restored permissions.
  • Improves subagent environment startup, provider-catalog correctness, Windows sandboxing, and SQLite reliability.
  • Makes queued input recover after reconnection without duplicate submissions.

Claude Code v2.1.287

Migration recommended
Fri, Oct 2, 2026Claude Code GitHub Releases
  • Introduces Claude Mods, enabling plugins to modify deeper Claude Code behavior.
  • Adds a built-in side-agent mod that flags missed considerations in telemetry-enabled first-party sessions.
  • Defaults supported cloud-provider deployments to a 1M context window for newer Opus and Fable models.
  • Changes MCP URL-prompt handling; servers failing after upgrade may require `bareElicitationCapability: true`.

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

New capability
Fri, Oct 2, 2026Hugging Face Blog
  • Releases an open MoE training stack behind Allen AI’s next-generation Olmo models.
  • Redesign targets efficient trillion-parameter-scale training with improved throughput over its earlier FSDP approach.
  • Exposes infrastructure for developers researching MoE routing, parallelism, and hardware adaptation.
  • Signals more capable open Olmo models with larger datasets and longer context windows ahead.

How to Gate Pull Requests on LLM Evals in CI

Nice-to-have
Fri, Oct 2, 2026OpenRouter Blog
  • Shows how to version agent eval cases beside prompts and agent code.
  • Uses a GitHub Actions status check to block merges when measured eval pass rates regress.
  • Recommends job-level path filtering so required checks do not remain pending or silently bypass prompt changes.
  • Covers repeated sampling, thresholds derived from baseline variance, and separating provider failures from regressions.

Claude Code v2.1.286

Migration recommended
Thu, Oct 1, 2026Claude Code GitHub Releases
  • Fixes session recovery after parallel tool-call crashes and large cloud-session wake failures.
  • Hardens MCP authentication, secret masking, connector discovery, and usage attribution.
  • Adds fallback to the previous same-tier model when the configured model is rejected.
  • Enforces disabled Remote Control policy changes by disconnecting active sessions.

OpenCode v1.18.34

Migration recommended
Thu, Oct 1, 2026OpenCode GitHub Releases
  • Adds namespaced session and parent-session identity headers to model requests.
  • Re-signs locally compiled macOS binaries for reliable execution on macOS 27 and later.
  • Signs macOS CLI release binaries with a Developer ID.
  • Fixes cross-platform plugin-name extraction in the status dialog.

Codex CLI 0.159.3

Nice-to-have
  • Adds optional account-security setup reminders for eligible local sessions signed in with ChatGPT.
  • Keeps the reminder opt-in rather than interrupting normal coding workflows.
  • Delivers a focused patch release with no agent-runtime or API migration.

Disrupting a coordinated model-distillation campaign

New capability
Thu, Oct 1, 2026OpenAI Blog
  • OpenAI disrupted a coordinated campaign that extracted protected reasoning from its models, with activity traced from early July and high-volume spikes in late July.
  • The operators did not break encryption or access stored conversations; they manipulated model interactions at scale. OpenAI attributes a core cluster to individuals associated with Moonshot AI.
  • Mitigations strengthened hidden-reasoning protections across users, workspaces, and model families, and added checks on streamed output.
  • Systems that support portable or replayable reasoning artifacts may face related risks; review tool-output and cross-session isolation assumptions.

Codex CLI 0.159.0

New capability
Wed, Sep 30, 2026OpenAI Codex GitHub Releases
  • Opt-in `instant_interrupt` lets fresh input steer Codex during a response or long-running code-mode call.
  • New TUI welcome, headers, warning handling, transcript scrolling, and expanded Mermaid rendering improve interactive agent use.
  • Windows MCP/code-mode launches, proxy-enabled sandbox networking, command-denial preservation, and `.aws` directory protections received reliability and safety fixes.
  • The old automatic follow-up suggestions setting was removed.

Claude Code v2.1.285

New capability
  • `claude --desktop` opens Claude Desktop on the current directory or an existing session, tightening terminal-to-desktop workflow handoff.
  • Plugin configuration can now be inspected and supplied at installation time, including bundled MCP-server settings.
  • `allowedProviders` gives administrators a managed allowlist for Anthropic API, Bedrock, Vertex AI, Foundry, gateways, and other supported providers.
  • Adds a WebFetch kill switch and fixes for subagents, permissions, plugins, secrets redaction, remote sessions, and long-running commands.

Anatomy of a bug-fixing agent

Nice-to-have
Wed, Sep 30, 2026Hugging Face Blog
  • Hugging Face details Serge, a nightly CI agent that identifies persistent Transformers failures, reproduces them on GPUs, patches them, verifies results, and opens reviewable PRs.
  • The agent runs in isolated short-lived pods with restricted network access and can only push `serge/` branches, keeping merge authority with maintainers.
  • Its verification funnel rejects flaky, duplicate, or weak patches before PR creation; 29 fixes landed over roughly 80 days.
  • Repository-history retrieval and repeated GPU tests show a practical pattern for unattended, safety-bounded maintenance agents.

Introducing GPT-6.1 Sol

New capability
Wed, Sep 30, 2026OpenAI Blog
  • GPT-6.1 Sol targets complex coding, computer use, and professional work, approaching GPT-6 Astra-level intelligence at one-fifth of its standard token prices.
  • Standard API pricing is $2 per 1M input tokens and $10 per 1M output, with cached input at $0.10 per 1M and a separate tier above 272K input tokens.
  • Adds multi-agent delegation in beta through the Responses API; tool-calling workloads must use the Responses API.
  • Available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu accounts; it is not yet available in Chat.

DevDay 2026 Recap

Nice-to-have
Wed, Sep 30, 2026OpenAI Blog
  • Official recap of the 20+ DevDay 2026 announcements across models, ChatGPT, Codex, APIs, security, and developer tools.
  • Serves as the index page for everything shipped at the event, including GPT-6.1 Sol and the new Agents API capabilities.
  • A useful starting point before reading the individual release posts.

Computer use added to the Agents API

New capability
Wed, Sep 30, 2026OpenAI API Changelog
  • Agents can now complete tasks in an OpenAI-hosted browser through the Agents API.
  • Website access approvals and user sign-in remain under application control.
  • Removes browser-execution infrastructure from many agent prototypes.
  • Treat browser authorization and account-session handling as first-class production concerns.