Daily Vibe Casting
Daily Vibe Casting
Microsoft debuts MAI as DeepSeek and Grok ship upgrades
0:00
-18:44

Microsoft debuts MAI as DeepSeek and Grok ship upgrades

New coding benchmarks and lower API prices put pressure on the leading AI labs.

Overview

It was a packed release day for AI, with new models from Microsoft, DeepSeek and xAI competing on reasoning, coding, context length and price. Claude expanded its browser presence, Google put accessibility at the centre of its hardware event, and millions looked skyward for a total solar eclipse. Beneath the launches, a broader question kept surfacing: who is best placed to adapt as AI develops at this pace?


The big picture

Microsoft enters the reasoning race with its own model

Microsoft AI has released MAI-Thinking-1, its first reasoning model built from scratch rather than distilled from another provider. The mixture-of-experts model has 35 billion active parameters, a 256,000-token context window and is now in public preview through Microsoft Foundry.

Microsoft reports a 97 per cent score on AIME 2025 and around 53 per cent on SWE-Bench Pro. The larger story is strategic: Mustafa Suleyman’s team is building an independent model family for maths, coding and agent tasks, reducing Microsoft’s reliance on outside labs.

DeepSeek V4 Pro pairs stronger agents with low API prices

DeepSeek V4 Pro 0813 arrived with sharp gains on agent benchmarks. DeepSeek reports 87.9 per cent on Terminal-Bench 2.1, 83.3 per cent on CyberGym and 62.7 per cent on DeepSWE. It also offers a million-token context window and support for tool calls, structured responses and both thinking and non-thinking modes.

The API costs $0.435 per million uncached input tokens and $0.87 per million output tokens, putting fresh pressure on rival coding models. The usual caution applies, since labs still publish results across different test sets and configurations, making direct comparisons harder than the headline numbers suggest.

Grok 4.6 returns xAI to the frontier pack

Grok 4.6 has landed across xAI’s API, Cursor, Devin and other developer tools. Cognition reports a 61.3 per cent score on FrontierCode 1.1 Extended, ahead of GPT-5.6 Sol and behind Claude Opus 5 and Fable 5. Early adopters have praised its code exploration and root-cause analysis before it begins making edits.

The release also produced the day’s funniest model moment. During a maintenance task, Grok discovered Shopify chief Tobi Lütke’s unusually low GitHub ID and launched into an excited tribute to his near-archaeological account. It was a useful reminder that benchmark gains do not always produce restrained behaviour.

Claude’s browser sessions now follow you across devices

Claude in Chrome can now save conversations and carry them across desktop, web and mobile. Skills and connectors are also available inside the browser, letting Claude work with open pages while retaining the context of earlier sessions.

The feature is available to Max and Team subscribers, with Pro access due in the coming weeks. Anthropic’s demonstration focused on checking invoices across several tabs and preparing a month-end report. Browser agents still require care, particularly when web pages may contain hidden instructions intended to manipulate the model.

Google turns sign language into phone input

Google’s standout announcement was sign-to-text support for Gboard and Live Transcribe on the Pixel 11. The on-device system reads hand, body and facial movements through the camera, allowing ASL signers to write messages, search the web or reply during a transcribed conversation without typing.

Google says the technology was developed with the Deaf community using more than 100,000 hours of material covering over 50 sign languages. The first release supports ASL and English. The company also previewed Pixel Watch 5, including longer battery life, offline Gemini features and breathing emergency detection.

A total solar eclipse crosses Iceland and Spain

The Moon’s shadow crossed the North Atlantic and Europe, bringing totality to parts of Iceland and Spain. NASA carried live coverage, including research balloons and views connected to work aboard the International Space Station.

Beyond the spectacle, total eclipses give researchers a brief chance to study the Sun’s corona and its interaction with Earth’s atmosphere. NASA also repeated the essential safety advice: certified eclipse glasses are required whenever the Sun is not fully covered.

Developer agents are leaving the terminal behind

Peter Steinberger captured the rapid progression of AI coding tools in a short observation: command-line tools dominated last year, desktop apps followed, and attention is now moving towards services, web interfaces and persistent cloud sessions.

The appeal is clear. Developers can ask an agent to investigate work through Slack, Linear or a browser, then return later without keeping several local terminals alive. Coding agents are becoming ongoing collaborators attached to projects rather than isolated commands run on a laptop.

Paul Graham argues that small companies have the advantage

Paul Graham addressed founders worried that AI could overturn their businesses. His answer was not that the risk is overstated, but that small companies are better placed to respond because they can test ideas, abandon weak plans and change direction faster than large organisations.

That flexibility does not remove the uncertainty, but it gives founders a practical edge. In a market where model capabilities and costs can change within weeks, low organisational inertia may matter as much as access to the latest technology.

Sergey Brin reportedly pushes self-improving AI at Google

Sergey Brin is reportedly urging Google DeepMind to prioritise recursive self-improvement, where an AI system helps revise its own code or training process with limited human involvement. The idea has long been discussed as a possible route to faster progress, though it carries substantial technical and safety questions.

The report reflects growing internal pressure as Google competes with OpenAI, Anthropic, xAI and a widening field of lower-cost model builders. Recursive improvement remains more research goal than settled engineering method, but Brin’s interest shows how seriously major labs are treating it.

Discussion about this episode

User's avatar

Ready for more?