Overview
Cheaper compute ran through the day’s biggest stories. Berkeley showed huge models running at useful speeds on consumer hardware, Grok topped a coding benchmark at a low cost per task, and OpenAI cut GPT-5.6 Sol prices. Elsewhere, Harvey and Anthropic pushed models deeper into legal work and security, while developers found new ways to make websites and coding tools easier for agents to use.
The big picture
FreeToken puts huge models on modest hardware
UC Berkeley’s Sky Lab has open-sourced FreeToken, a serving system that moves work between GPU memory, system RAM and the CPU as conditions change. The reported results are striking: Qwen3.6-35B reached 39.3 tokens per second on an 8GB RTX 4060 laptop, while a single RTX PRO 6000 ran the 753B-parameter GLM-5.2 at nearly 15 tokens per second.
The team reports speeds two to four times higher than Ollama on consumer GPUs. If those gains hold across wider testing, capable local inference will no longer require a workstation packed with premium cards.
Grok takes the coding crown as model prices fall
Grok 4.6 Extra High reached 70.8 per cent on CursorBench 3.2 at $2.81 per task. That narrowly beat Fable 5 Max and Opus 5 Max, while costing far less to run. Benchmarks remain snapshots, but cost per completed task is becoming as important as the raw score for coding agents that may work through long jobs.
OpenAI added to the price pressure by cutting GPT-5.6 Sol API rates by more than 20 per cent for three months. Together, the announcements point towards tougher competition around sustained agent work, where small differences in price compound quickly.
Harvey trains an open model for legal work
Legal AI company Harvey has released Tenet, its first model post-trained specifically for legal tasks. It starts with Moonshot AI’s open-weight Kimi K3 and reportedly nearly doubles the base model’s completion rate on the LAB legal agent benchmark.
Harvey says Tenet also performs strongly on contract evaluations while running at less than a quarter of the cost of leading foundation models. The broader lesson is that firms with deep subject knowledge and good evaluation data can build strong specialist systems without training a foundation model from scratch.
Claude Mythos 5 starts hunting security flaws
Claude Security scans now use Mythos 5 in a public beta for Claude Enterprise customers. The service examines GitHub repositories, follows data flows across files and returns possible vulnerabilities with severity ratings, confidence scores, CWE categories and suggested repairs.
This goes beyond spotting suspicious lines in isolation. Repository-wide reasoning may help uncover faults that only appear when several components interact, though security teams will still need to verify findings before changing production code.
Websites are being graded for agent access
Vercel has launched is-agentic.com, a tool that checks whether automated agents can discover, read and use a website. Its audits include more than 100 checks, visual records of an agent’s journey, suggested fixes and a command-line interface.
The release arrives as developers debate how bots should identify themselves online. A separate popular post encouraged agents to imitate approved crawler identities to bypass blocks, exposing how weak header-based access rules can be. Site owners now face two linked jobs: making public material legible to legitimate agents while enforcing boundaries that cannot be defeated by changing a name in a request.
Ox Alpha arrives with vast free capacity and a mystery owner
Nous Research has opened limited free access to Ox Alpha and claims capacity for a quadrillion tokens per day. The anonymous model offers a million-token context window, multimodal input and tool use, with an apparent focus on coding and long-running agent tasks.
Researchers examining its tokenizer and API behaviour believe it belongs to Z.ai’s GLM-5.X family, possibly GLM-5.3. That attribution remains unofficial, but the launch shows how anonymous model trials can create interest before a formal release or even a confirmed developer.
Anthropic takes a step towards its own chips
Anthropic has hired Amir Salek, who previously led Google’s custom TPU programme, as it explores an in-house semiconductor business. The company currently buys compute from Nvidia, Google and Amazon to train and run Claude.
Designing chips would be expensive and slow, but it could give Anthropic tighter control over cost, supply and the hardware choices behind its models. The hire also shows that competition between AI labs is extending below the model layer and into the physical infrastructure.
Claude Code gets steadier away from the desk
Anthropic has responded to complaints about Claude Code’s Remote Control feature with automatic reconnection, live syncing between phone and laptop, clearer device cards and faster handling of heavy sessions on iOS. Sessions can now be started from a phone and followed across devices with fewer interruptions.
Another popular internal practice surfaced alongside the update: Anthropic staff use an ELI5 command that asks Claude to explain unfamiliar subjects through simple HTML pages with large diagrams and little text. Both ideas reflect the same goal, making agents easier to direct and understand during ordinary work.
Useful products are outrunning polished names
Daniel’s joke that “branding does not matter” landed because ChatGPT is such an unadorned name. It combines a plain activity with a technical acronym, yet became a household term because people immediately understood the product and found it useful.
Martin Casado made a related argument about AI investment: relatively small amounts of money can now produce a measurable result quickly, rather than funding years of engineering before anyone knows whether the idea works. In that environment, product feedback arrives early, and clever naming has less room to rescue a weak experience.
Falcon 9 sends another 29 Starlink satellites to orbit
SpaceX launched 29 Starlink satellites from Cape Canaveral on the Falcon 9 booster B1078, which was flying for the 30th time. The first stage returned to the droneship A Shortfall of Gravitas after an earlier launch attempt was scrubbed.
The striking part is how routine such a technically demanding mission now appears. Reusing the same booster 30 times while maintaining a regular launch schedule is a reminder that repetition, not spectacle, is what turns ambitious engineering into infrastructure.
























