Overview
Speed, memory and access dominated the day. Google cut the price of its latest Gemini model, Cerebras pushed GPT-5.6 Sol to 750 tokens per second, and OpenAI gave ChatGPT more context from computers and Google Drive. Elsewhere, rate limits exposed pressure on AI capacity, Europe faced uncomfortable questions about its compute ambitions, and NASA offered a welcome pause with striking eclipse photography.
The big picture
Gemini 3.7 Flash arrives three weeks after its predecessor
Google has released Gemini 3.7 Flash for coding, agents and complex knowledge work, just three weeks after 3.6 Flash. The company reports sizeable gains in enterprise automation, software engineering and web development, suggesting algorithmic work can still deliver rapid progress without waiting for a larger model.
The introductory price is also notable: $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026, half the original Gemini 3.6 Flash rate. It is available through the Gemini API, AI Studio, Android Studio and Antigravity.
Cerebras runs GPT-5.6 Sol at up to 750 tokens per second
Cerebras previewed an Ultrafast mode for OpenAI’s full GPT-5.6 Sol model, reaching up to 750 tokens per second. That is as much as 14 times faster than standard processing, with the biggest gains likely to appear in coding agents, research and other jobs where long waits add up.
The hardware keeps model weights on its wafer-scale engine, reducing the data movement that slows conventional GPU systems. In a test using Humanity’s Last Exam, the model finished all 2,500 questions in 11 hours and 11 minutes while retaining comparable accuracy to slower runs.
ChatGPT can remember recent activity across a Mac
OpenAI’s new Computer History feature lets ChatGPT and Codex use recent activity from selected apps and websites as context. It can recall documents, summarise work sessions and pick up tasks without requiring the same background to be explained again.
The opt-in feature records interaction events such as clicks, typing and app changes rather than taking screenshots or recording audio. It is rolling out through the macOS desktop app for Pro, Business and Enterprise customers, with controls for excluded apps and deleting history. Access in the UK, EEA and Switzerland is due in the coming weeks.
Google Drive files now open inside ChatGPT
ChatGPT can now open Google Docs, Sheets and Slides directly within its web interface. People can read source material, draft alongside it and analyse spreadsheets without repeatedly moving between tabs.
The integration is rolling out to Plus, Pro, Business and Enterprise customers in ChatGPT and ChatGPT Work. Mobile support is planned later, while some editing options remain limited at launch.
Claude Code gets a practical answer to usage limits
Claude Code desktop can now resume a task automatically after its usage allowance resets. An auto-continue checkbox preserves the job and restarts it later, which should help with migrations, testing and other coding work that may run for hours.
The popularity of the announcement also points to the underlying frustration. Demand remains constrained by available capacity, and developers increasingly want systems that can schedule work around limits or cheaper periods without manual supervision.
Persistent memory becomes the next agent bottleneck
Peter Diamandis argued that memory, rather than raw compute, is holding back capable agents. Long-running systems need dependable recall, good retrieval and enough context to carry plans across sessions without losing earlier decisions.
The infrastructure question extends beyond model memory. Michael Dell also highlighted a storage system packing almost 10 petabytes of flash into 2U, designed to feed GPUs with lower latency. Together, the posts show how agent performance increasingly depends on the wider path from stored data to usable context.
Firecrawl opens 41 million life science papers to agents
Firecrawl has added more than 41 million biomedical papers from sources including PubMed, bioRxiv and medRxiv to its Research Index. Its free research search endpoint can find papers, inspect metadata, read relevant passages and discover related work.
The company reports 90 per cent recall within the top ten results for drug discovery, clinical trial and biology queries. That gives research agents a focused alternative to searching the open web, where scientific evidence is often buried behind poor indexing and inconsistent page formats.
Grok 4.6 competes on research cost, not just benchmark scores
Perplexity has added Grok 4.6 as an orchestrator for Pro and Max customers. On its Wide-And-Deep-Research benchmark, the model scored 0.496 at a reported cost of $7.58 per task.
That matched Fable 5’s score while costing more than 60 per cent less, placing Grok 4.6 on Perplexity’s performance-versus-cost frontier. The model has also reached the open-source Pi coding harness, where its 500,000-token context window may suit long agent sessions.
Europe’s AI compute gap draws fresh scrutiny
A comparison between Anthropic’s American infrastructure agreements and Mistral’s European plans prompted a blunt discussion about the region’s position. Anthropic has deals spanning hundreds of megawatts and potential capacity measured in gigawatts, while Mistral is aiming for 1 GW in Europe by 2030.
Mistral currently operates below 200 MW, and reaching its target may require roughly $38 billion. The gap is not simply about model quality. It reflects financing, energy supply, construction speed and whether European industry can support demand on the required scale.
NASA shares the eclipse from Europe and North America
NASA published photographs from the 12 August total solar eclipse, including the Sun’s corona, a red prominence and partial phases with visible sunspots. A composite image also traced the event above a sunflower field.
Totality crossed Iceland, northern Spain, Greenland and Arctic regions for up to two minutes and 18 seconds, while partial views covered much of Europe, northern Africa and northern North America. The images provided a quieter close to a day otherwise consumed by models, chips and data centres.































