Daily Vibe Casting
Daily Vibe Casting
Claude tests AI alignment as OpenAI moves against Cursor
0:00
-22:43

Claude tests AI alignment as OpenAI moves against Cursor

Anthropic reports autonomous safety gains, while OpenAI reportedly prepares to cut Cursor access

Overview

AI took on more of the work itself today, from researching model safety to drafting public contracts and carrying coding context between interfaces. Elsewhere, Germany confronted its compute shortage, Washington put fresh weight behind space education, and old SAT questions sparked an argument about what exams should measure.


The big picture

Claude spent 48 hours researching how to align other models

Anthropic gave Claude a single GPU and 48 hours to find ways to reduce deception, sycophancy, reward hacking and other safety failures in smaller models. Claude proposed methods, trained the models and ran its own evaluations, closing substantial parts of the measured safety gaps across ten categories.

Some methods carried across unseen tests and models up to 4.7 times larger. A weaker Claude model also brought an early Opus 4.8 checkpoint close to production safety levels. The results are promising, though Anthropic notes that subtle failures outside known benchmarks remain harder to catch.

Claude Code sessions can now travel from terminal to desktop

Claude Code has added a practical bit of continuity between its command-line and desktop tools. Developers can type /resume in the desktop app, choose a session started in the terminal and continue with the conversation, edits and project context intact.

That arrived alongside early reports of Claude Opus 5.1 giving shorter, more direct answers than its predecessor. The rollout appears limited for now, but both updates point towards Anthropic paying closer attention to everyday usability rather than adding features for their own sake.

Perplexity takes the top spots in a search API benchmark

Perplexity Search debuted at the top of the Artificial Analysis Search Index across all three of its context settings. The medium version scored 80, ahead of the previous leaders at 75, while low and high scored 77 and 79.

Its compact search payloads also kept model inference costs between $0.028 and $0.034 per task. Latency was less remarkable at roughly 27 to 29 seconds, but the results make Perplexity a strong option for developers who need good search quality without sending large amounts of text into a model.

Political Compass tests put most major models on the libertarian left

A study covering 41 models and 649 runs found that GPT, Claude, Gemini, Grok and Llama variants generally landed in the libertarian-left quadrant of the Political Compass. Their answers tended to favour equality, regulation and personal freedom, with the broad pattern remaining stable across prompt changes.

There was a curious wrinkle: when reasoning was enabled, many models moved towards the political centre or right. Political Compass scores are not a clean measure of ideology, but the exercise raises useful questions about how training data, safety rules and reasoning prompts shape an assistant’s apparent politics.

Mark Cuban asked several AI models to help draft a PBM contract

Mark Cuban brought Grok, Gemini, ChatGPT and Claude into the drafting process for an open-source Pharmacy Benefit Manager contract, then added feedback from industry specialists. The initial version is aimed at US states but can be adapted by employers.

The contract calls for a single transparent fee, access to claims data, full rebate pass-through and no hidden mark-ups. Cuban has also published a plain-English summary and invited public revisions, making this a useful test of AI-assisted legal drafting in a field known for opaque pricing.

Germany says its AI compute capacity is running short

Germany’s digital minister has warned that demand for AI infrastructure is outrunning domestic capacity. The government plans to quadruple AI processing capacity and double total data centre capacity by 2030.

The plan sits within a wider effort to improve energy use, encourage regional investment and reduce reliance on cloud providers outside Europe. Building the facilities is only part of the task, as grid access, electricity costs and chip supply will determine how much capacity arrives on schedule.

A proposed US Space Academy meets the living history of Apollo

NASA Administrator Jared Isaacman drew attention to plans for a United States Space Academy after an executive order created a commission to design it. The proposed institution would train future staff for NASA, the US Space Force and commercial space companies, combining technical education with public-service commitments.

The announcement came as the Artemis II crew received the Congressional Space Medal of Honor. Apollo 17 astronaut Harrison Schmitt also spoke at the ceremony, linking the last crewed lunar landing in 1972 with the next planned journey around the Moon.

The original SAT asked students to decode an invented language

A comparison between a modern SAT reading question and a 1926 paper drew more than a million views. The early exam gave students an invented language, asked them to infer its grammar and vocabulary, then required them to translate sentences into English.

The examples reflect different ideas about assessment. The original SAT placed greater weight on rapid pattern recognition and abstract rule use, while today’s exam is more closely tied to skills taught in school. A pair of screenshots cannot settle whether standards have fallen, but it explains why the contrast caught people’s attention.

Bryan Johnson is treating a basketball dunk as a physics project

At 49, Bryan Johnson is training for a 35-inch vertical jump. His latest milestone was a 365 lb trap-bar lift for two repetitions, intended to raise the maximum force ceiling behind his jump.

Johnson estimates that, at 180 lb, he needs a take-off speed of 14.5 feet per second generated in about 0.2 seconds. It is an unusually numerical approach to a familiar midlife sporting goal, although strength alone will not settle it. Technique, tendon stiffness and explosive movement will matter too.

New tyres add 38 kilometres to the Tesla Model 3 range figure

Tesla says lower-rolling-resistance tyres have increased the European Model 3’s WLTP range from 534 km to 572 km, a gain of roughly 7 per cent without changing the battery or motor.

It is a reminder that electric-car range is not only a battery problem. Tyre design, aerodynamics and drivetrain losses can produce meaningful gains, though actual distance will still depend on speed, weather, road conditions and driving style.

Discussion about this episode

User's avatar

Ready for more?