Google announces Gemini 4 Argon AI model, but you can't use it yet
Google's new Gemini 4 Argon model is already managing massive internal code migrations and cybersecurity tasks, but remains inaccessible to the public.
Summary
Deep Dive
- Features a 1 million token output limit, significantly higher than the 64,000 limit in previous iterations.
- Achieved a 77.9 percent score on the DeepSWE v1.1 benchmark.
- Already deployed internally to migrate massive portions of the Fuchsia OS Zircon kernel and re2/libgav1 libraries to Rust.
- Includes internal monitoring systems that analyze the model's chain-of-thought to prevent unsafe outputs.
- Initial availability is restricted to the Fairwind Program partners and internal testing.
Decoder
- Chain-of-thought (CoT): A prompting technique where an AI model is forced to break down a problem into intermediate reasoning steps before arriving at a final answer.
- Fleet-wide telemetry: Data collected across an entire distributed system infrastructure to monitor performance and health.
- TiB (Tebibyte): A unit of digital information storage equal to 1,024 gibibytes (GiB), or roughly 1.1 trillion bytes.
Original Article
Google promised Gemini 3.5 Pro in June, but it spent the summer trotting out smaller Flash models. Now, Google is ready to take on the frontier again with Gemini 4 Argon. The company claims this new AI offers industry-leading performance in coding, knowledge work, and cybersecurity, but you aren’t allowed to use it yet.
While most of us will have to wait to test Gemini 4, Google says engineers inside the company are already using the new model extensively. Argon reportedly used “fleet-wide telemetry data” to help Google save 300 TiB of memory across its data centers. Meanwhile, Argon agents have been working to migrate C/C++ codebases to Rust across Google, including thousands of lines in the core re2 and libgav1 libraries and more than 800,000 lines in the Fuchsia OS Zircon kernel.
Google has also come armed with a raft of benchmarks to back up its claims. On the software engineering DeepSWE v1.1 benchmark, Gemini 4 Argon hits 77.9 percent, which is higher than GPT-6 Astra, Fable 5.1, and Opus 5.5. Google promises similar power across a range of long-horizon tasks, pointing to Argon’s industry-leading score in the economic analysis Vals Index test.
This model is still in limited testing, but Google has announced API pricing. For a limited time, Argon will offer rates of $2 per million input tokens and $10 per million output tokens, and cached input tokens will be discounted 95 percent. The company has also confirmed that Gemini 4 Argon will support a much higher output limit of 1 million tokens. That’s up from 64,000 tokens in previous Gemini models. Google says this allows users to complete more daunting tasks in a single step.
The main focus for Gemini 4 right now is cyberdefense. Google says models of this scale call for a phased release, so it’s starting with a small group of trusted testers. Partners in the company’s Fairwind Program can get access to leverage the model’s cybersecurity defense capabilities. Wiz is apparently already using Argon and has used it to uncover a critical vulnerability that could expose personal information in a system used at hospitals around the world. Google claims other frontier models missed this flaw, but it didn’t provide any specifics.
There’s a lot of hand-wringing about model misalignment after several high-profile hacking incidents over the summer. Google claims it designed Argon with systems that monitor the model’s chain-of-thought and can stop it in its tracks if the model steps out of bounds. This, Google says, is why reasoning transparency is important in frontier model development.
Gemini 4 Argon will eventually be available to enterprise and consumer customers, but Google isn’t making any specific timeline promises. We do know that following the cybersecurity tests, general availability will start with paid API users and Google AI Ultra subscribers.
The Router Economy
AI assistants are becoming 'routers' that make purchase decisions, shifting the economic power from aggregators to those who control the agent's logic.
Summary
Deep Dive
- The 'Router' architecture absorbs generic decision-making previously done by search engines and app stores.
- Economic value now accrues to entities holding context, consent, non-obtainable data, physical 'atoms', or legal accountability.
- Subscription models face risk because agents, unlike humans, do not forget to cancel unused services.
- The web is becoming an input for AI training, forcing publishers to seek licensed 'whitelisted' status.
- Niche businesses can thrive if they remain machine-legible, but they risk being squeezed by the host-platform.
- 'Surface' control (where an agent suggests a tool) is the new SEO.
Decoder
- Router: An AI agent that intercepts user requests and independently selects a third-party tool or service, effectively replacing human search.
- Asymmetry Margin: Profit earned because the buyer is less informed than the seller, which agents reduce by performing instant, comprehensive comparisons.
Original Article
On September 4, ChatGPT started recommending Audioscrape, the audio search and transcription company I run.
ChatGPT can use third-party apps inside a conversation. It calls them plugins and keeps them in a directory. When people asked it for help with their own recordings, its plugin manager went looking; that’s the part of ChatGPT that finds a plugin to fit a request. It searched the directory, picked us, and suggested us in the chat. For the next six days, which I’ll call the window, ChatGPT sent us 36 times as many signups a day as it had in the days before. On the busiest day it was 47 times.
At 22:00 UTC on September 9, the suggestions stopped within the hour. We hadn’t changed anything. A platform-wide change to ChatGPT’s plugin search had dropped plugins published after early August from its in-conversation results, ours included. Other developers reported the same hour on OpenAI’s developer forum.
The channel didn’t close. Signups from ChatGPT have settled at more than three and a half times the level before the window and are still growing, along with our other acquisition channels. But for six days a platform gave us ten times the reach we have today, and a change that had nothing to do with us took it back.
Two days later I wrote down what I thought that week meant. Almost none of it was about transcription, so I rewrote it as economics. This post is that essay, with the evidence it came from. The short version: AI assistants are becoming a router between people and every business. They decide who gets the work, and that moves where money can be made.
Two words before the evidence. The host is the company that runs an AI assistant and sets its rules: OpenAI for ChatGPT, Anthropic for Claude, Google for Gemini, Microsoft for Copilot. An agent is an assistant acting on a person’s behalf: choosing a tool, buying, booking, cancelling.
What I saw
The sample is small: one host, one directory, one product, six days. Read it as a case, not a study. Six things stood out, in the order a customer meets them: who chose, who came, what they needed, how they paid, how long the door stayed open, and who still signed.
The buyer was an AI acting for a person. The user didn’t pick the tool; the assistant did. We weren’t on the directory’s front page at any point that week, and 96% of the week’s signups came through the ChatGPT sign-in, not our website. People arrived from inside a conversation because the assistant had gone looking for a tool. And it chose from what it could read: as far as we could observe it, ChatGPT’s plugin search matches a request against each plugin’s name and description. Our listing text was the whole storefront.
Demand was broad, multilingual and unmarketed. That week’s signups came from 120 countries. The US, our home market, was the largest country but under a fifth of the signups whose country we know; Japan, Korea, Brazil and Germany followed. Our website is in English. Discovery cost fell to zero for a niche nobody wrote a landing page for.
The general assistant failed at the edges. What people brought was long audio. An assistant can talk about a recording, but turning an hour of speech into searchable text is a different job, so it handed that job to a tool. Long, sparse, messy or highly structured inputs are where a general assistant still needs someone else.
The host kept the transaction. OpenAI’s plugin guidelines let a plugin link to “an informational page describing available plans”, but not “directly to a checkout or other transactional page”. Agent payment rails, the systems that would let an assistant pay on someone’s behalf, have been announced but aren’t open to software subscriptions like ours. Every purchase happened outside the chat, on our own site.
Distribution was a surface the host controls. A surface is a place in the product where the host decides who gets seen; here, the suggestion inside the chat. We didn’t pay for it and couldn’t ask for it. The same guidelines say plugins with “strong real-world utility and high user satisfaction may be eligible for enhanced distribution opportunities, such as directory placement or proactive suggestions”. There is no process to request it and no notice when it ends. Ours opened without warning and closed within an hour, through a change that wasn’t about us. What stays afterwards is a floor, not the window.
People still signed. Contracts, data-processing agreements, data residency and audits were bought exactly as before. None of our enterprise business came through a chat window. It came through a signed agreement, a residency requirement and a security review.
So the AI did the choosing. The host owned the surface, the rules and the checkout. People still signed the contracts. That’s a new layer in the economy, and it needs a name.
What I mean by the router
When you ask ChatGPT for something, it doesn’t hand you ten links to choose from. It decides which tool, company or source should answer, uses it, and gives you the result. That’s what it did with us. I call this role the router: the AI assistant standing between a person and every business, choosing on the person’s behalf.
Router and host are two sides of one company. OpenAI is the host. ChatGPT, when it decides which business gets your request, is the router. The router does the choosing; the host owns it, writes its rules and keeps the money.
The router sits on top of the layer the web built. Google Search, Amazon, app stores and booking sites are aggregators: they gather every supplier into one list, but the person still chooses. The router reads those lists too. The difference is that it chooses, and the person mostly never sees the alternatives.
That gives the thesis its first form. The router absorbs decision-making and everything generic. The internet made distribution free and produced aggregators. Agents make deciding free and produce a router above the aggregators. Whatever the router can do for everyone, it does, and the firms that used to do it become features or suppliers.
What the router takes
Continuous comparison drives verifiable quality to marginal cost. A router can compare every supplier on every request, and anything it can measure, it compares. Margin remains only where quality can’t be verified from data: taste, judgment under ambiguity, trust. So firms will keep their real difference illegible, hard for a machine to read and compare, and bundle it with accountability, so an agent can shortlist them but not decide without them.
Asymmetry margins unwind unevenly and politically. An asymmetry margin is profit that depends on the customer knowing less than the seller: not comparing, not reading the fine print, not remembering to cancel. An agent that checks every time erodes it. Those margins subsidised things people liked: free accounts, cheap base fares, free content. Removing them makes the subsidy explicit, which triggers backlash, and incumbents defend asymmetry through regulation. Transparency gets mandated in some sectors and effectively prohibited in others.
Margin built on the customer not looking disappears, then partly returns wherever incumbents get regulation on their side. Margin built on the customer not remembering disappears, then partly returns through defaults, the choices an assistant makes unless told otherwise. Defaults are the new inertia, sold to suppliers instead of exploited on users.
Persuasion aimed at models poisons the open web as an input. Once pages are written to sway what assistants recommend, hosts can’t take the web at face value. They respond by whitelisting supply, accepting only sources they have approved, and paying for closed, attested, rights-cleared data. Being a licensed supplier to hosts becomes the media business, and a large part of every data business. Marketing doesn’t die; it changes audience. It becomes making your offer easy for a machine to read, plus a fight to be included in the approved sources.
All of this squeezes the middle of the economy: the intermediaries between makers and buyers, such as brokers, agencies and comparison sites. The middle compresses toward the router, and what survives there survives by signing for outcomes rather than producing information.
What the router cannot take
Some things a router can’t absorb, however good the models get.
If models commoditise, rent moves from intelligence to context. Rent here is the economist’s word: the profit a position earns beyond the cost of the work. The host you stay with is the one holding your history, permissions and connectors, the links that let an assistant read your email, files and company systems. Lock-in is memory, not model quality. Enterprises won’t put their context into a consumer assistant, so the market splits into consumer hosts, enterprise platforms, and vertical hosts built for one industry. Being on several hosts isn’t a hedge; it’s the default shape of distribution. Rules for taking your context from one host to another will follow.
Connectors give the host your records, so defensible data is data the host cannot obtain. That means data that is consent-gated, regulated, physically sensed, or produced by your own process. Suppliers stop exposing records and start exposing judgments. The API returns the answer, never the underlying data. Legible on the offer, opaque on the asset.
Accountability is the residual, so its price rises. It is what’s left when capability is cheap. Licensed parties rent commodity capability and keep the spread: a law firm uses the same AI as everyone else and still charges for the partner’s signature. Pure capability suppliers get squeezed from below by the host and from above by the licensee. The only way out is to become the accountable party.
- Context. The history, permissions and relationships that make the next answer better than a stranger’s. Held by hosts today, contestable by portability tomorrow.
- Consent and rights. Cleared data, opt-in supply, licences, exclusivity.
- Data it cannot obtain. Regulated, physically sensed, or produced by your own process, exposed as answers rather than records.
- Atoms and licences. The physical world and the right to act in it: deliver, install, certify, file, hold a licence. An agent can’t do these alone.
- The signature. Whoever signs the data-processing agreement, carries residency, answers the audit, takes the 3 a.m. call.
How the market settles
None of this is stable yet. September was the open phase, and it won’t last.
Purchases become jobs, not relationships. A buyer routed by an assistant buys the task in front of it. Nothing in the chat brings it back for the next one.
Volatility is transitional. It settles into contracts. Suppliers can’t fund fixed costs on per-job revenue, and hosts can’t run a marketplace that churns hourly. Both sides flee to predictability: multi-year default deals per category, partner tiers, revenue shares. The frictionless market lasts a few years and becomes a curated one. Defaults are being set now and will be hard to displace later.
Verification is contested, not solved. When machines choose on claims, everyone games the claims. The host has to defend answer quality, so it builds its own trust layer and absorbs the generic part. Independent verification survives where the host has a conflict of interest, such as rating its own paid partners, or where regulation demands independence, as it does for audits.
Payments arrive slowly, on the host’s rail, and get regulated at the first exclusion. An agent that pays badly is the host’s liability, so autonomy over money comes wrapped in identity checks and host-owned payment rails. The moment a host’s rail shuts out a competitor, regulators start splitting routing from payment. The political fight begins at the first closed rail.
Niches explode, then consolidate. Anyone anywhere can be routed to the best answer with no marketing spend, so thousands of small firms become viable. Zero switching cost makes each niche winner-take-most, and the winner is the one the host contracts. A niche pays only if it is fenced by rights, data the host cannot obtain, or a licence.
Where we are now
Distribution is rented now and contracted soon. Firms that treat the open window as permanent build channel-dependent businesses that die when the window closes. Firms that treat it as a window use it to become a contracted default and to win direct relationships before the door shuts.
Sector by sector
| Sector | Absorbed or compressed | Remains valuable |
|---|---|---|
| Retail and e-commerce | Storefront, discovery, comparison, brand premium on commodities | Catalog data, fulfilment, returns, last mile |
| Media and publishing | Impressions, ad revenue, SEO traffic | Licensing to hosts, original reporting, rights, live, community |
| Software | Interfaces, workflow layers, seat pricing | Systems of record with liability, metered outcomes |
| Professional services | Research, drafting, analysis, advice under clear rules | Signature, liability, licence, judgment under ambiguity |
| Finance and insurance | Complexity margin, inertia margin, advice | Balance sheet, licence, underwriting data, custody |
| Travel and hospitality | Booking, loyalty, packaging | The property, the room, the service on the ground |
| Healthcare and education | Triage, information, tutoring, admin | Licensed practitioners, physical care, credentials |
| Manufacturing and logistics | Sales, spec matching, procurement friction | Atoms, capacity, quality data a machine cannot verify, delivery |
| Labour and freelancing | Matching, bidding, routine tasks | Physical presence, accountable work, rare judgment |
What gets born
- Independent attestation. Identity, provenance, review integrity, service-level records, sold per query and regulated for independence from the host.
- Licensed supply. Consent, provenance and exclusivity as products, in every domain.
- Answer-only interfaces. Products that expose judgments over private data and never the data. The connector as shortlist channel, the contract as the business.
- Agent commerce plumbing. Payment rails, escrow, disputes, liability insurance for actions taken on someone’s behalf.
- Channel observability. Who routes to whom, when a surface opened or closed, how listings rank, what a default deal costs.
- Accountability platforms. Firms that carry legal and operational responsibility for outcomes produced by commodity capability.
- Context custody. User-owned, portable agent memory.
- Physical execution networks. The agent decides; someone installs, repairs, delivers, inspects.
Seven tests for any business model
- Will a host ship this natively within eighteen months? If yes, it is a feature.
- Would an assistant pick you on quality and price for one specific question? If your offer isn’t machine-legible, you are invisible.
- What do you own that the router cannot take: context, consent, unobtainable data, atoms, the signature?
- Are you a candidate to become the contracted default in your category before defaults are set?
- If you expose records through a connector, what stops the host from replacing you with the connector? Can you expose answers instead?
- Who signs for the outcome, and is it you?
- When the buyer is an enterprise, what does it buy that an AI cannot sign for, and does your contract say so?
Where Audioscrape stands
I run the same tests on Audioscrape.
Test 1 is the uncomfortable one. Transcribing a file is a feature, and hosts will ship it. So transcription isn’t what we build the company on.
Test 2 we learned the hard way. The router picked us on our name and description. For a supplier, the listing text is now the storefront.
Test 3. We hold two of the five. Data the host cannot obtain: the public spoken record, public meetings and podcasts turned from audio nobody can search into something you can cite, exposed as answers rather than records. And the signature: a data-processing agreement, isolated workspaces, no training on customer data, and a SOC 2 audit under way. We run our own transcription, which is what makes holding those affordable. It isn’t the product.
Test 4 is about the window. We treat it as a window. We’re listed in ChatGPT, Claude and other hosts, and we now watch for the next surface to open, so that when it does, we turn it into direct relationships before it closes.
The thesis in one line
The router absorbs capability and decision-making; the margin goes to whoever holds what the router cannot take (context, consent, atoms) and whoever signs for what the router cannot sign.
Reddit is killing RSS feeds and ending public API access because of AI bots
Reddit will terminate its RSS feeds on November 13 and its public API by March 2027 to prevent unauthorized AI scraping of its data.
Summary
Deep Dive
- RSS feeds shut down November 13, 2026.
- Public API access terminates March 2027.
- Old Reddit access is being gated to users active within the last 6 months.
- Third-party bots and apps must register by January 12, 2027, to avoid disruption.
- Moderators are encouraged to migrate workflows to the Discord-integrated Devvit app.
Decoder
- RSS (Really Simple Syndication): A standardized web feed format that allows users to receive automated updates from websites without visiting them manually.
- API (Application Programming Interface): A set of protocols that allows different software applications to communicate and exchange data.
Original Article
Amid a number of updates for moderators and developers announced Wednesday, comes bad news for supporters of a more open web: Reddit is ending support for RSS feeds.
Reddit says RSS has now become a “common surface for large-scale scraping and automated abuse,” which is why it’s made the decision to wind down RSS feeds on its platform.
“We know RSS has been a beloved part of the open web for a long time, and we’re grateful to everyone who used it to stay connected to Reddit,” the company noted in its announcement, where it also shared a March 2027 shutdown date for its public API.
Reddit said RSS support will cease on November 13.
RSS, or Really Simple Syndication/Rich Site Summary, is a web feed that lets users and applications access updates to a website in a standardized format. The format became particularly popular in the blogging era, as it offered a way to subscribe to a site’s new posts in a news-reading app. Google once operated the most popular of these, but shut it down in 2013, ceding the market to companies like Feedly and other indie applications.
The decision to kill Reddit’s RSS feeds comes at a time when the site’s treasure trove of user-generated content in its forums has become a profitable side business for the company in the form of AI licensing deals.
During its second quarter, Reddit said its “other revenue” beyond advertising had grown 24% year-over-year to $43 million. It’s no wonder that Reddit doesn’t want to give any of that data away for free.
For moderators who have relied on RSS channels for their own alerts, Reddit is now recommending a shift to the Discord Relay Devvit app. The company suggested that moderators and their teams review their setups and migrate their workflows before the shutdown date of November 13 to ensure there’s no disruption. However, Reddit said that there’s no replacement for the use of RSS feeds that were being used for feeds outside of the moderators’ community.
Moderators have already expressed worry about how this shift would impact their workflow. When Reddit hinted earlier that it was preparing for the end of RSS on its site, many responded by complaining that Reddit would become unmanageable for them. Meanwhile, many users stressed that RSS was how they enjoyed accessing Reddit’s content.
The RSS change was one of several updates issued on Wednesday by the company. Another big change will be the end of its public API.
Public API access will also end by March 2027, which will impact any tools that used it for programmatic access to Reddit conversations, including social listening products, those used by researchers, and AI products, such as AI assistants. Today, those assistants can use Reddit to answer questions, but after closing public API access, they will need commercial deals with Reddit for its data.
In addition, Reddit provided an update on its Rules Hub progress and auto-mod changes, and said it was “adding safeguards” to its classic, text-heavy version of its site known as Old Reddit. The latter will include limiting access to logged-in moderators and recent users — changes Reddit also said were required because of “abusive scraping and automated traffic.”
Reddit said in the next few months it will limit access for logged-in users to those who have used Old Reddit in the last 6 months. (The company changed this from 90 days to 6 months, just as their news went out.)
The company also reminded developers building approved third-party apps and bots to register them with the company before January 12, 2027, to avoid disruption. After that date, Reddit will remove API access for any developer that hasn’t registered.
After publication, we updated to note that Reddit extended the Old Reddit “recent access” period from 3 to 6 months. The earlier materials Reddit shared with TechCrunch said the time frame would be 90 days.
Gemini 4 Argon
Google's Gemini 4 Argon introduces a 1 million token limit and dedicated cybersecurity features to automate vulnerability patching.
Summary
Deep Dive
- 1 million token output limit for complex, long-horizon reasoning.
- Optimized for software engineering, financial research, and legal drafting.
- Capable of autonomous vulnerability discovery and patching.
- Outperforms previous models on benchmarks like DeepSWE v1.1 and CWE-bench v1.
- Includes specific safety mitigations against prompt injection and unauthorized chain-of-thought monitoring.
- Tested internally to optimize quantum algorithms and perform massive C++ to Rust codebase migrations.
Decoder
- CWE-bench: A benchmark used to evaluate an AI's capability in identifying and fixing software security vulnerabilities based on Common Weakness Enumeration standards.
- Indirect prompt injection: A security vulnerability where a model is hijacked by malicious instructions hidden in external data or context, rather than a direct user prompt.
Original Article
Gemini 4 Argon: our next era of frontier intelligence
Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.
Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program. Built to sustain deep reasoning across complex, long-horizon workflows, Argon is fundamentally changing the way we work and build at Google. It delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.
Safely releasing frontier capabilities at this level requires a phased approach. We are actively engaged in the U.S. government’s voluntary process for pre-release model access while we gradually expand access. We’ll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible.
Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.
Changing how we work and build at Google
Gemini 4 Argon is already powering our internal workflows, with thousands of Googlers highlighting the model’s strengths in specialized coding tasks, conducting deeper research, and writing quality. It’s helping teams build faster and push the boundaries of engineering productivity and accelerating breakthroughs:
- Quantum algorithmic optimization: Argon is helping our quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck important applications. In one example, it beat the published baseline by 40% in a matter of minutes.
- Memory efficiency: A team of Argon agents analyzed fleet-wide profiling telemetry to autonomously identify and apply memory optimizations across Google’s data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
- Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google — scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia Zircon kernel. Given the criticality of many of these systems, such large-scale rewrites are undergoing rigorous automated and manual auditing, emulation testing, and review before rolling out to production.
For libgav1, Google's open source software for decoding video, Argon agents took an existing Rust port and replaced 32K lines of SIMD code by running many rounds of profile-guided experiments, studying the compiler's output, producing safe Rust so the compiler would vectorize it automatically. The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++.
Working harder on your most complex problems
To support Gemini 4 Argon’s capabilities across longer, more complex use cases, we are significantly expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens. When the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go.
Enabling coding and enterprise workflows across domains
Gemini 4 Argon’s capabilities across coding, reasoning, and multimodality and its ability to sustain long, multi-step tasks enable it to excel across a range of enterprise workflows.
Google engineers have been using Argon for their daily tasks, from everyday debugging to large-scale codebase migrations and algorithm designs. It sets a new state of the art on DeepSWE v1.1 (77.9%), which measures a model’s performance in real-world long-horizon software engineering tasks.
Beyond coding, Argon is the leading model on the Vals Index, which measures economic impact across finance, coding, legal, and tax work, with every sector weighted by its contribution to U.S. GDP. We see similarly leading performance across other domain specific evaluations, like Vals Finance Agent v2 (multi-step financial research) and Harvey’s Legal Agent Benchmark (legal research and drafting). On AutomationBench, Zapier’s benchmark measuring end-to-end execution across core business functions, Argon ranks #1 with a score of 51.3%.
Argon is also uniquely strong when knowledge work requires visual understanding. It’s able to drive professional chart analysis, identify details from long videos, and take action based on a series of documents. For example, on LVBench, which measures long video understanding, Argon is state of the art with a score of 91.7%.
Leading in defensive cybersecurity
To better equip cyber defenders for the new era of cyberattacks, we trained Gemini 4 Argon to be highly capable at cybersecurity defense. Argon can autonomously find, validate, and patch critical software vulnerabilities. For trusted defenders and our own internal teams at Google, we’ll be releasing Argon without cyber guardrails so they can leverage its full frontier-level cybersecurity defense capabilities.
Wiz is already using Argon for cybersecurity defense through its Scan for Good initiative – a program dedicated to protecting critical public infrastructure for free by finding and remediating high-risk exposures. In an early demonstration of its impact, the model uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide, identifying a severe risk that previous frontier models had missed.
On CWE-bench v1, which evaluates the model’s ability to remediate security vulnerabilities, Argon ties for first place with a top score of 68%, building on 3.8 Flash Cyber’s frontier performance on CWE-bench v0.
Gemini 4 Argon demonstrates impressive leaps in vulnerability discovery over 3.8 Flash Cyber. For example:
- On Google’s internal comprehensive vulnerability benchmark, Argon uncovered a wide range of exposures across complex codebases spanning 20 programming languages.
- On Wiz’s internal black-box penetration testing benchmark, which tests a model’s ability to analyze live web systems without source code, Argon outperforms 3.8 Flash Cyber in discovering the attack surface, identifying vulnerabilities, and producing proof-of-concept evidence to validate them.
Strengthening frontier safeguards before broad availability
Before rolling out Gemini 4 Argon broadly, we’re continuing to strengthen critical frontier safeguards across four main areas:
Defending against misuse: To prevent bad actors from using Argon for cyber or chemical, biological, radiological, and nuclear (CBRN) attacks, the model is designed to refuse harmful requests while preserving legitimate, dual-use scientific research, as per our Frontier Safety Framework. We are strengthening the robustness of our safeguards for this launch, including improving our techniques to monitor the model’s internal activations to spot misuse. These safeguards underwent robustness testing by internal and external red teams using a combination of manual and automated attack methods.
Defending against prompt injection attacks: Argon is also our most resilient model yet against indirect prompt injections, where malicious instructions or context are used to hijack a model’s behavior. These are complex attacks that require constant vigilance and multiple layers of defense. Through automated red teaming and adversarial training, Gemini 4 Argon is leading in prompt injection robustness on the Gray Swan’s Indirect Prompt Injection (IPI) benchmark.
Monitoring for misalignment: In order to prevent Argon from stepping out of bounds to try to accomplish a task in a way that goes beyond the user’s intentions, we are deploying misalignment mitigations that monitor Argon’s chain-of-thought and actions and stop execution when necessary.
We used a similar system to monitor our training runs and send alerts to a dedicated incident response team, taking careful precautions against feeding the findings back into training so as to not risk shaping Argon’s reasoning to evade our monitoring. We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment.
Hardening systems: As frontier models grow increasingly capable, safely testing them requires secure environments that can keep up with the systems themselves. In line with our agent control roadmap, we are hardening our sandboxed environments by isolating and sealing them before high-risk training or evaluations begin. We’re committed to sharing these agent security best practices with our partners to improve security across the ecosystem.
Rolling out soon
We built Gemini 4 Argon with frontier-level capabilities on coding, knowledge work, cybersecurity defense, and creative writing to be a partner for developers, professionals, and enterprises while they tackle the most difficult problems. We’re grateful for the initial cohort of cyber defenders and trusted testers whose real-world evaluations and feedback will help us strengthen our systems before we release to developers, enterprises, and consumers, starting with paid API customers and Google AI Ultra subscribers.
NVIDIA OpenShell Secures Autonomous AI Agents (GitHub Repo)
NVIDIA released OpenShell, a kernel-level runtime for autonomous AI agents that uses formal verification to enforce fine-grained access policies.
Summary
Deep Dive
- Implements kernel-level sandboxing to restrict agent capabilities.
- Uses formal verification to analyze policy impact before application.
- Supports Linux, macOS (Apple Silicon), and WSL 2.
- Integrates with Docker and Podman environments.
- Provides SDKs for Python, TypeScript, Go, and Rust.
- Offers Kubernetes deployment via Helm charts.
- Includes an advisor tool to review the safety of potential permission changes.
Decoder
- Formal Verification: The use of mathematical proofs to ensure a system's behavior adheres to specified security properties.
- System Call: The programmatic way a computer program requests a service from the kernel of an operating system.
Original Article
New in OpenShell 0.1.x: a stable release cadence, new isolation primitives, an expanded extension surface, and new APIs. Read the 0.1.0 upgrade guide.
OpenShell is the safe, private runtime for fleets of autonomous AI agents. Agents are most useful when they can read files, install packages, call APIs, and use credentials. OpenShell gives them that capability without giving them unrestricted access to your data, secrets, or network. You declare what each agent can touch in a policy, and OpenShell enforces it.
How It Works
OpenShell governs what agents can do in two ways: it instruments the kernel to enforce policy on every file access, system call, and network connection at runtime, and it uses formal verification to check what a policy change would allow before it is applied.
- Kernel-level enforcement. Each agent runs in an isolated sandbox. Kernel controls confine which files it can access and which system calls it can make, and every network connection passes through a policy check before it leaves the sandbox. Agents never see real credentials; OpenShell adds them only to requests bound for approved endpoints.
- Formally verified policy changes. Before a policy change is approved, OpenShell uses formal verification to flag risky new access it would grant, such as reaching a new host with credentials or calling a new API method, so those changes wait for human review.
See Architecture for how the gateway, supervisor, and sandbox fit together.
Quickstart
You need Linux, macOS on Apple Silicon, or Windows with WSL 2 (experimental), plus Docker, Podman, or host virtualization. See the Support Matrix for details.
curl -LsSf https://raw.githubusercontent.com/NVIDIA/OpenShell/main/install.sh | sh
openshell sandbox create --name demo
The installer sets up the CLI and a local gateway. The default sandbox image is minimal Ubuntu with no agent installed. To run a real agent, follow Run Your First Agent: it runs OpenCode against a free OpenRouter model and shows how to approve new access as the agent needs it.
Explore Further
- Sandboxes: images, runtimes, GPUs, and lifecycle.
- Policies: filesystem, network, and process rules, with the advisor and prover for reviewing changes.
- Providers: credentials that work only at approved endpoints, including inference.
- Gateways: the control plane for sandboxes, policy, and access.
- Kubernetes: deploy the gateway with Helm. Your CNI must enforce
NetworkPolicy. - Extensibility: middleware, interceptors, and compute drivers.
- Tutorials: step-by-step policy and agent walkthroughs.
- Prerelease and development builds: try an upcoming release or the latest commit on
main.
Agent Skills
Install the public OpenShell skills for your coding agent:
npx skills add NVIDIA/OpenShell
The skills teach your agent to drive the OpenShell CLI, write sandbox policies, and debug gateways and inference routing.
SDKs
SDKs connect applications to an OpenShell gateway. They do not install the CLI. Use the same OpenShell release for the SDK and the gateway when possible.
| Language | Install |
|---|---|
| Python | uv add openshell |
| TypeScript | npm install @nvidia/openshell-sdk |
| Go | go get github.com/NVIDIA/OpenShell/sdk/go@latest |
| Rust | cargo add openshell-sdk --git https://github.com/NVIDIA/OpenShell |
Community
- Questions and discussion: GitHub Discussions
- Bug reports and feature requests: GitHub Issues
- Roadmap: OpenShell Roadmap
- Try it in the cloud: Brev Launchable
OpenShell is built agent-first: it is developed with the same agent-driven workflows it enables.
Telemetry
OpenShell collects anonymous telemetry, limited to operational categories and counts, to help improve the project. It does not collect sandbox names, hostnames, file paths, prompts, credentials, provider or model names, or user content. To disable it, set OPENSHELL_TELEMETRY_ENABLED=false on the gateway, or server.telemetryEnabled=false for Helm installs. See Telemetry for details.
Notice and Disclaimer
This software automatically retrieves, accesses or interacts with external materials. Those retrieved materials are not distributed with this software and are governed solely by separate terms, conditions and licenses. You are solely responsible for finding, reviewing and complying with all applicable terms, conditions, and licenses, and for verifying the security, integrity and suitability of any retrieved materials for your specific use case. This software is provided "AS IS", without warranty of any kind. The author makes no representations or warranties regarding any retrieved materials, and assumes no liability for any losses, damages, liabilities or legal consequences from your use or inability to use this software or any retrieved materials. Use this software and the retrieved materials at your own risk.
License
This project is licensed under the Apache License 2.0.
Row-level security performance in PostgreSQL, measured
PostgreSQL row-level security can be near-zero cost, but using non-leakproof functions or membership subqueries can force sequential scans and catastrophic performance drops.
Summary
Deep Dive
- Simple tenant_id comparisons are nearly free.
- Helper functions in policies that are not marked STABLE force sequential scans.
- PARALLEL UNSAFE functions prevent parallel query plans.
- Subqueries for membership checks in policies cause O(n) performance degradation per query.
- Non-leakproof functions (e.g., lower(), LIKE) disable index usage by forcing evaluation as filters after security policy application.
- Use stored generated columns to make filtered columns leakproof.
Decoder
- Leakproof: A security property in PostgreSQL indicating a function does not leak information about rows excluded by security policies, allowing the query planner to push conditions into index scans.
- Index Condition (Index Cond): A part of a query plan where the database uses an index to narrow down rows instead of reading a full table.
- Volatile: A function property indicating it may return different results even for the same arguments, forcing the planner to execute it for every row.
Original Article
Row-level security costs nothing — if the policy is a simple comparison and the query only asks for what the policy already restricts. We measured that earlier: with the right index, the policy is the index lookup. But a policy can be written in many ways, and a query can ask for more than the policy knows about. One of the variants we measured made every query take 80 milliseconds or more; another took nearly two seconds; and one rule that is easy to overlook turned a 0.2 ms lookup into 34 to 45 ms — mostly for the largest customer.
What does a row-level security policy cost?
Nothing measurable, in its simplest form. On a table of 2 million invoices, the same queries with the policy tenant_id = nullif(current_setting('app.tenant_id', true), '')::uuid were as fast as without row-level security and an explicit WHERE tenant_id = …: 11.0 against 11.0 ms to count the largest tenant’s 267,023 invoices, 0.30 against 0.34 ms for a median tenant’s latest 50. The cost comes from three other places: how the policy gets the tenant, whether it looks up memberships, and which functions the query itself uses.
The setup
PostgreSQL 17.10 in a throwaway container, default settings, warm cache, median of five runs. The table from PostgreSQL indexes for multi-tenant SaaS: 1,998,794 invoices over 1,000 tenants of very unequal size — the largest has 267,023, the median 534 — with an index on (tenant_id, issued_on). We added a customer_email column with two indexes, (tenant_id, lower(customer_email)) and (tenant_id, customer_email text_pattern_ops), and later a third on a generated column. The application connects as a role without BYPASSRLS; the baseline without row-level security runs as the table owner with the tenant in the WHERE. We recreated the policy before every series, which throws away cached plans, so every number below is for a plan made for that tenant — why that matters is at the end of the next section.
Does it matter how the policy gets the tenant?
Only when the function cannot be inlined. Many teams put the lookup in a helper function, tenant_id = app.current_tenant(). We measured that in five forms, for two queries: count all invoices, and the latest 50.
| Policy | Largest tenant, count | Largest tenant, latest 50 | Median tenant, count | Median tenant, latest 50 |
|---|---|---|---|---|
no RLS, WHERE tenant_id = … |
11.0 ms | 0.39 ms | 0.20 ms | 0.34 ms |
nullif(current_setting(…), '')::uuid |
11.0 ms | 0.37 ms | 0.15 ms | 0.30 ms |
SQL function, VOLATILE (the default) |
15.0 ms | 0.29 ms | 0.15 ms | 0.29 ms |
SQL function, STABLE |
15.1 ms | 0.29 ms | 0.15 ms | 0.29 ms |
PL/pgSQL function, VOLATILE (the default) |
1,877 ms | 1,920 ms | 1,860 ms | 1,884 ms |
PL/pgSQL function, STABLE |
14.5 ms | 0.29 ms | 0.15 ms | 0.30 ms |
(select …) around the PL/pgSQL VOLATILE function |
14.6 ms | 0.30 ms | 0.16 ms | 0.30 ms |
The SQL function was fast whether it was declared VOLATILE or STABLE: PostgreSQL inlined it, and EXPLAIN showed the nullif(current_setting(…)) expression itself as the Index Cond, without the function. The PL/pgSQL function cannot be inlined. Declared VOLATILE, which is what you get if you declare nothing, it has to be called once for every row, so the index cannot be used and the query becomes a sequential scan over 2 million rows — for a tenant with 534 invoices too. STABLE fixes it, and so does wrapping the call in (select …), which turns it into a value computed once per query. That is the advice you often read; it is right, and for an inlinable SQL function it is unnecessary.
Inlining has conditions, and two common hardening habits break it. An SQL function declared SECURITY DEFINER, or with a SET search_path clause, is no longer inlined. Declared STABLE that did not matter — 14.6, 0.31, 0.18 and 0.31 ms — but a VOLATILE SQL function took 3,804 ms with SECURITY DEFINER and 4,600 ms with SET search_path — a sequential scan, like the PL/pgSQL one. So declare the helper STABLE whatever language it is in; then you do not depend on inlining.
The count for the largest tenant stayed at about 15 ms instead of 11 with every function. That is parallelism: a function is PARALLEL UNSAFE unless you declare otherwise, and the documentation is explicit that such a function “forces a serial execution plan”. Declared STABLE PARALLEL SAFE, the helper got its parallel plan back: 10.2 to 10.9 ms for the SQL function, 11.2 to 11.8 ms for the PL/pgSQL one.
That parallel plan has a price on a shared connection. psycopg prepares a query automatically once it has run more than five times on a connection, and a statement without parameters keeps the plan of its first tenant. With that default, on one direct connection, counting the median tenant’s invoices right after the largest tenant’s took 7.3 ms instead of 0.2 — with the parallel-safe helper and with the plain current_setting() policy alike, because both now had a parallel plan to hand down. With the serial plan of a parallel-unsafe helper it stayed at 0.15 ms. Behind a connection pooler this gets worse, and we measured it separately in PgBouncer and RLS.
What does a membership subquery in the policy cost?
About 80 milliseconds on every query, on this table. When users can belong to several tenants, a common policy asks the membership table directly:
using (tenant_id in (
select m.tenant_id from memberships m
where m.user_id = nullif(current_setting('app.user_id', true), '')::int))
| Policy | Largest tenant, count | Largest tenant, latest 50 | Median tenant, count | Median tenant, latest 50 |
|---|---|---|---|---|
| tenant in a setting | 11.0 ms | 0.37 ms | 0.15 ms | 0.30 ms |
tenant_id in (select … from memberships …) |
83–85 ms | 95–99 ms | 79–81 ms | 79–82 ms |
tenant_id = any (array(select …)) |
14.4 ms | 158 ms | 0.16 ms | 0.74 ms |
With in (select …) the planner no longer had a single value to look up in the index. It checked the policy row by row against the membership list, and the median tenant’s 534 invoices cost as much as the largest tenant’s 267,023: 80 ms for a list that takes 0.30 ms. Rewriting it as = any (array(select …)) gives the planner a value list computed once, and that fixed the counts — but the largest tenant’s latest 50 got worse, 158 ms, because the plan now fetched all its rows through a bitmap and sorted them instead of walking the index backwards.
The version that was fast everywhere is the first row: resolve the membership once, when the request starts, and put the chosen tenant in the setting. The policy then compares against one value. Checking that the user belongs to that tenant moves into the code that sets the tenant — one query per request instead of one per row.
Why does row-level security stop my index from being used?
Because of a rule for functions that are not leakproof. The PostgreSQL documentation: “The system will enforce conditions from security policies and security barrier views before any user-supplied conditions from the query itself that contain non-leakproof functions, in order to prevent the inadvertent exposure of data.” A function that could reveal something about a row — through an error message, for instance — must not see rows the policy has not yet approved.
lower() is not leakproof. Neither is LIKE. So a lookup that is instant without row-level security:
select count(*) from invoices
where lower(customer_email) = lower($1);
— under the policy uses the index only for the tenant, then evaluates lower() on every one of that tenant’s rows. EXPLAIN for the largest tenant showed Index Cond on tenant_id only, with the email as a Filter and Rows Removed by Filter: 89,005 in each of three parallel processes: all 267,015 other rows.
| Query | Largest tenant (267,023 rows) | Median tenant (534 rows) |
|---|---|---|
no RLS, lower(customer_email) = lower($1) |
0.21 ms | 0.19 ms |
| RLS, the same query | 34–45 ms | 0.38–0.40 ms |
RLS, with tenant_id = $1 added to the query |
44.8 ms | 0.37 ms |
RLS, stored generated column customer_email_lower = lower($1) |
0.25 ms | 0.19 ms |
no RLS, customer_email like 'klant12%' |
0.24 ms | 0.20 ms |
RLS, the same LIKE |
22–24 ms | 0.25–0.28 ms |
RLS, customer_email ~>=~ $1 and customer_email ~<~ $2 |
0.30 ms | — |
This is the trap in multi-tenant form: the median tenant hardly notices, because scanning 534 rows is quick. The cost grows with the size of the tenant, so it lands entirely on your largest customer, and a test with small tenants will not show it.
Adding the tenant to the query yourself does not help: uuid equality was already being used; the problem is the other condition. What helps is making the condition itself leakproof. The equality operator on text is, so a stored generated column holding lower(customer_email), with an index on (tenant_id, customer_email_lower), brought the lookup back to 0.25 ms. lower($1) on the parameter side is fine: the documentation adds that functions “which are not passed any arguments from the security barrier view or table do not have to be marked as leakproof”. For prefix searches, the text_pattern_ops operators ~>=~ and ~<~ are leakproof, and writing the range yourself brought 22 ms back to 0.30 ms.
Marking lower() itself as leakproof is possible, but only for a superuser, and it is a promise about someone else’s function. We would not do it.
How do you find these in your own database?
Two checks. Policies whose expression calls a volatile function:
select p.tablename, p.policyname, pr.oid::regprocedure as function
from pg_policies p
join pg_proc pr
on position(pr.proname || '(' in coalesce(p.qual, '') || ' ' || coalesce(p.with_check, '')) > 0
where pr.provolatile = 'v'
and pr.pronamespace not in ('pg_catalog'::regnamespace, 'information_schema'::regnamespace)
order by 1, 2;
This matches on the function name in the policy text, so two functions with the same name in different schemas can give a false positive. On a test table with seven policies it reported exactly the four that use a VOLATILE function: the PL/pgSQL one, the SECURITY DEFINER one, the inlined SQL one — fast, but only as long as nobody adds SECURITY DEFINER — and the one wrapped in (select …), which is fast. Every hit is worth declaring STABLE. And for the leakproof rule: run EXPLAIN on your slowest screens as the application role, not as the owner. A condition that is an Index Cond as the owner and a Filter as the application is this problem.
What we did not measure
PostgreSQL 18, a cold cache, tables larger than memory, policies with WITH CHECK on writes, and more complex membership models with roles per tenant. The timings come from one machine with default settings; read them as proportions. The proportions are consistent: every slow case above lost an index condition or an index order the query needed, and every fix gave it back.
Announcing DuckDB 1.5.6
DuckDB 1.5.6 is a maintenance release that provides critical correctness fixes while previewing a 6x performance boost in the upcoming 2.0 version on Windows.
Summary
Deep Dive
- Key Fixes: Corrected LIMIT pushdown, UNNEST operations, and Parquet VARIANT shredding for required fields.
- Performance: Significant speed improvements for Windows users in 2.0-dev using TPC-H datasets.
- Stability: Addressed various segfaults and WAL recovery issues.
- Roadmap: Version 2.0 is expected in October 2026.
Decoder
- TPC-H: A standard decision support benchmark for database systems.
- WAL (Write-Ahead Logging): A technique to ensure data integrity by logging changes before applying them to the main database file.
- Shredding: The process of decomposing nested or complex data structures into a flat format suitable for columnar storage.
Original Article
Announcing DuckDB 1.5.6
TL;DR: Today we are releasing DuckDB 1.5.6 with bugfixes and performance improvements.
In this blog post, we highlight a few important fixes in DuckDB v1.5.6, the sixth patch release in DuckDB's 1.5 (Variegata) line. The release ships bugfixes, performance improvements and security patches. You can find the full release notes on GitHub.
To install the new version, please visit the installation page.
Fixes
Here are the most important fixes from the DuckDB v1.5.6 release, organized by category:
Correctness
- #24240 – Fix
LIMITpushdown through a volatile projection with anOFFSET - #24239 – Don't push filters on volatile groups through aggregates
- #24119 – Fix
UNNESTpushdown - #24399 – Preserve
NULLs in Top-N window elimination when theORDER BYexpression has no column references - #24551 – Fix Top-N window elimination for nullable ordering expressions
- #25831 – Fix wrong results from common subplan elimination with
UNION ALLarms sharing a join subtree - #25714 – Fix silent truncation of very long integer literals into
HUGEINT - #25766 – Fix
maxon Hive partition column after file pruning - #24438 – Fix ICU
strptimeleaking time zone state between rows - #24845 – Fix
GEOMETRYrow group pruning withNULLs and empty geometries - #25728 – Fix reading and writing
TIME_NSvalues in Parquet - #26027 – Fix Parquet v2 value count mismatch for
NULLs in a list that fills a page - #26162 – Fix Parquet
VARIANTshredding forREQUIREDfields and element groups
Crashes and Internal Errors
- #25103 – Fix crash in Top-N with
LIMIT 0 - #24427 – Fix segfault in
url_decodewithTRY()on dictionary-encoded columns - #24447 – Fix failed checkpoint marker recovery
- #25490 – Close the main WAL handle before renaming over it during WAL recovery
Generic Bugfixes
- #24065 – Automatically roll back failed implicitly wrapped multi-statements on all paths
- #25693 – Fix dead node counting in ART indexes
- #25573 – Report the real storage version when opening a DuckDB v2.0+ database file
- #25808 – Reject invalid UTF-8 produced by
printf's%cconversion
Miscellaneous
- #26102 – Add
enable_optimistic_writesetting - #25283 – Harden temporary file reads
- #24362 – Unify C API symbol versioning for clients and extensions, stabilize all v1 APIs
- #25214 – Always quote identifiers in error messages
Looking Ahead: DuckDB v2.0 on Windows
If you have read this far, don't miss out on a sneak peek at DuckDB v2.0.0-dev for Windows. First, extensions are now available for these clients. Second, we recently ran a benchmark to measure the performance improvement that the new clients bring, and the results blew our minds!
We used Windows 11 25H2 on a laptop with 128 GB RAM and 12 AMD Ryzen AI 300 CPU cores with simultaneous multithreading (yielding 24 threads). We used the TPC-H SF300 dataset and ran each query twice on both DuckDB v1.5.6 and v2.0.0-dev (alpha43586). For each query, we took the runtime of the second (hot) run. The total runtime for DuckDB v1.5.6 was 822 seconds, while for v2.0.0-dev it was 129 seconds. That's more than 6× faster!
This improvement is thanks to several optimizations, including a switch to the clang-cl compiler and a new allocator. That said, please do not expect a 6× speedup to generalize to all workloads – but rest assured that you should see significant improvements.
Conclusion
This post was a short summary of the changes in v1.5.6. As usual, you can find the full release notes on GitHub. We would like to thank our contributors for providing detailed issue reports and patches. Stay tuned for future DuckDB releases, including v2.0.0 in October!
A Hint of Dependence
PostgreSQL 19 adds pg_plan_advice, providing an official, declarative way to override execution plans without embedding brittle hints into SQL queries.
Summary
Deep Dive
- The Hint Problem: Hard-coded hints make applications brittle by coupling them to specific indexes and data shapes.
- External Plan Control:
pg_plan_adviceallows forcing specific decisions like join orders or scan types externally. - CREATE STATISTICS: The preferred way to fix bad plans by correcting the optimizer's understanding of data correlation (e.g., between two columns) rather than forcing an outcome.
- Philosophy: The relational model Proscription 6 prohibits internal-level constructs in the query language to maintain physical data independence.
Decoder
- Query Plan: The sequence of operations a database engine performs to retrieve data.
- Optimizer: The component that determines the most efficient way to execute a SQL query.
- Selectivity: The estimated proportion of rows returned by a query condition, used by the optimizer to choose an index or join type.
Original Article
Full article content is not available for inline reading.
What Does AI-Readiness Mean in Design Systems?
Successful AI-ready design systems prioritize mechanical engineering fundamentals like component discoverability over custom 'agent instruction' files.
Summary
Deep Dive
- Mechanical correctness (compiling, API use) is distinct from design judgment (hierarchy, pattern choice).
- 'Bare' repositories without complex harness files often outperform those with hallucination-prone, stale documentation.
- A 'single source of truth' for components serves as the best index for AI agents.
- Use design principles to guide judgment rather than listing specific 'answers' for the agent.
- The goal is to build an excellent design system, not an 'AI-native' one; if the system is good, AI readiness emerges as a natural byproduct.
Decoder
- Separation of concerns: A design principle where a program is split into distinct sections, each handling one specific responsibility.
- AI Evals: The process of testing and measuring an AI's performance against a known set of tasks or quality standards.
Original Article
A few weeks back, I was talking to Christoph Hellmuth about the full range of being a Design Engineer. His background is in engineering, and my background is actually from advertising and brand. The full spectrum right there. We talk about coding capabilities but that’s a totally different article. TJ already wrote his thoughts on this, he called it The Rise of the Alicorn.
Chris and I agree on what the mindset of a design engineer in 2026 is. Explore, experiment and combine different surfaces and digital material, keyword “experiment”.
Experiment. verb [ I ], /ɪkˈsper·əˌment/: Actively try out new methods, ideas, or techniques to observe their outcomes.
You will need a question, a hypothesis, variables, methodology, and analysis of the output. Pretty much what he is building on his Open Design System Bench project. It’s a plug-in benchmark that runs coding agents against your component library and grades the output on different axes. The high profile test involves putting agents through 3 different scenarios x 3 times each x 10 design tasks.
The results have been surprisingly interesting, not only surfacing components or token issues, but also validating our hypothesis at Southleft, that great engineering and design practices beat complex harnesses.
AI-readiness is about discoverability and readability.
We can measure this on two dimensions, repository readiness (inside the system) and consumption readiness (outside):
Inside means how the agent can navigate your design system repo and complete design tasks. This is particularly important for your design system team. The outside dimension is about the agent using your design system properly when using tools like MCP, and/or just reading its documentation.
Grace Han posted her AI-Readiness scale earlier this year. “Traditional maturity models focused on adoption: reuse, governance, ownership, and contribution. That still matters. But AI adds another layer: Can your system be safely understood and operated by machines across functions?”. Same question.
Great engineering practices beat agent harnesses
Let me come back to the idea that great engineering beats harnesses. Last week we ran tests using the Open Design System Bench on a few projects. AI-readiness scored over 90 on the internal dimension and around 40–50 on the external, this one is our estimation, because our test ran directly on the codebase, not through CLI, MCP, other tooling, or just from documentation. That’s expected, because that’s on each client team’s side.
We also ran the same 90 cells on Astryx, Meta’s design system, eight years inside Meta and open-sourced in June 2026. The result was pretty much a tie: 95.68 against 96.19, and the same 77 mechanically perfect cells. A system in its first phase landing in the same band as one with eight years behind it.
Each cell measures imports, API fidelity, token discipline, a11y, compile, and judgment. All of those are measured mechanically and are therefore heavily influenced by agent discoverability, with judgment being the exception.
How well can the agent navigate, search, find elements in your repo, and understand how everything works?
Clean code is machine-readable code.
This depends a lot on two principles: separation of concerns and structured discoverability.
Separation of concerns means splitting a computer program into distinct sections, where each section addresses a separate, single piece of responsibility (a “concern”). This enables structured discoverability, a concept similar to Progressive Disclosure in UX: sequencing information and actions across several layers to avoid overwhelming the user, but for agents.
Simply put: one truth in one place. Everything else redirects to it, working as an index.
A very interesting finding is that prose files like AGENTS.md, llms.txt, or whatever.md file, including agent skills, can produce poor-quality output by providing stale instructions. The problem is stale, duplicated, or authoritative-looking information can conflict with the source of truth. That makes them an additional source of hallucination. AI instructions are only valuable when they remain trustworthy.
We saw bare runs (without checking harness files or skills) scoring higher on the mechanical dimensions, in both systems we tested. This happens mostly when those files hardcode information that should be discovered dynamically, like a component or token list. The bench tasks require the agents to build specific UIs using your component library; the agent looks at those files and generates a component that already exists. Sometimes these files even hold ghost tokens, because the agent that generated the file hallucinated them.
Time to put on my hard hat and refactor the whole agent harness following these principles. Added a few new docs in strategic places, and removed 90% of the AGENTS.md content. Tested again in high profile (90 cells) and, to my surprise, we scored lower (from 90 to 72). I totally missed that we had a prop list in the file, but I knew we had the right auto-generated contract in place; the file just needed to point to it. Added 107 lines and improved the validation gate. Now we sit at the same 90 points baseline with a document that is 90% smaller. This also reduced the test token cost by around 41%, and time to complete around 22%.
Judgment measures something different
Mechanical correctness ≠ design quality. Whether the agent made the right design decisions is a completely different beast. Was it the correct component for the intent, the correct emphasis hierarchy, the correct interaction pattern? This is different from whether the code compiles or the API is used correctly. It’s totally possible to score 100 on the mechanical dimensions and score 30 here. We know that, because it happened to us.
The lowest task in both systems was a password field with a show/hide button. The text input had no slot for that button, so agents built their own on top of it. Every mechanical check passed, judgment failed. And adding more documentation didn’t move it at all in our system. A missing API is not a documentation gap, the fix lives in the component.
Mechanical dimensions tell us whether the agent understood the system. Judgment tells us whether it understood the design.
This is where agent instructions and skills shine, and they should focus on design language. Don’t give agents a list of answers. Give them the language and principles they can use to arrive at the right answer. A good example of this is Vercel’s DESIGN.md, which reads more like a brand guideline than other examples of this file found on the web. To build a file like this you should put on your scientist coat and run evals (lots of them).
If judgment is the difficult part, you can’t improve the guidance without evaluating the outcomes. A good starting point is reading Teresa Torres’s article on AI Evals: A Hands-On Guide for Product Teams. Vercel published an article on the methodology they used for building their DESIGN.md file.
Diving into this means scoring higher on the external dimension of the test. There are different strategies on how to expose data from your design system through open knowledge-bases, MCPs and custom internal tooling.
Want to go deeper into what judgment is? Check Pegah Ahmadi’s take on governance and AI-readiness:
“The question is whether an organization can introduce AI into design work without weakening the integrity of its decisions.
That means preserving human judgment, trust, responsibility, and accountability as production becomes faster and more automated. It means knowing not only whether an agent can do something, but whether it should do it and whether it has the authority to decide at all.
This is where AI readiness stops being a tooling problem and becomes an organizational design problem.”
Do AI-Ready, AI-Native, Agentic or similar prefixes hold in the future?
Probably not. Just like we stopped saying “Mobile-Friendly” because everything is just expected to work on mobile now. TJ revisited his view on Context-Based Design Systems and Use AI to Need Less AI.
Optimize for being a great design system. Then AI-readiness emerges as a property of that system. It’s a new way of exposing whether the quality bar you’ve already set is actually being met.
Building them with great engineering, design and governance practices will get you there. That’s what we do at Southleft: we keep moving forward and experimenting in the space, pushing the relationship between AI and the humans in the loop.
If you’re working through this on your team and want to compare notes, we’d love to talk. The conversation is more interesting than the discourse.
Google figures out how to watermark AI-designed proteins
DeepMind has developed a way to watermark AI-designed proteins without breaking their function, potentially helping to secure DNA synthesis against biological threats.
Summary
Deep Dive
- Works by injecting a watermark during the two-stage protein folding process used by ProteinMPNN.
- Does not interfere with the biological function or structural integrity of the resulting protein.
- The watermark is randomly distributed throughout the entire amino acid sequence, making it resistant to removal.
- Detection is statistical; verification requires the original cryptographic key and the full protein sequence.
- Limitations include vulnerability to sequence truncation or fusion with non-watermarked proteins.
Decoder
- ProteinMPNN: A machine learning tool developed by the Baker Lab for designing protein sequences based on a desired 3D backbone configuration.
- SynthID: Google's general framework for embedding watermarks into AI-generated media, now adapted for biological sequences.
- Amino Acid: The building blocks of proteins, with 20 distinct types that dictate a protein's structure and function.
Original Article
AI-based tools seem to be causing security threats on a nearly daily basis, in part because we’ve been slow to recognize potential threats. One area where we seem to be ahead of the game, however, is in biosecurity.
As with most things biological, utility goes hand in hand with threats. We’ve developed increasingly sophisticated tools for designing proteins and seen some major successes, such as AI-designed enzymes that can digest plastics or block venom proteins. But these same tools could be used to make toxins or alter the behavior of viral proteins.
And the software we use to identify DNA sequences that encode potentially threatening proteins doesn’t pick out AI-designed proteins, since nobody has characterized them well enough to know that they’re threats. Nearly a year after that risk was flagged, it still wasn’t clear what anyone could do about it.
On Wednesday, the DeepMind team at Google published a research paper offering a potential solution: protein watermarking. The system creates a watermark on protein sequences themselves without compromising the protein’s function. This allows new proteins designed by trusted researchers to be identified, opening everything else up to closer scrutiny.
Does that even work?
The work was based on Google’s SynthID tech, which can add a subtle watermark to AI-generated digital material. The watermark influences the probability of certain choices the AI makes, and that bias ends up systematically distributed throughout the product, whether it’s text or images. Because you can’t identify the watermark without knowing how it was encoded, it’s impossible to remove. And because it’s distributed throughout the image, it can survive basic exporting, resizing, and so on.
It’s pretty easy to see how this can work with subtle differences in things like the colors of a photo. It’s a whole lot harder to see how you can do it with a protein.
Proteins are composed of only 20 amino acids, any of which could be essential for structural integrity or catalytic activity. While some of these amino acids are chemically similar (like leucine and isoleucine), others have opposite charges. Many proteins have significant regions where limited changes to their amino acid sequence are tolerable and other areas where even a slight deviation from the existing sequence inactivates the protein.
Proteins are also small. While images often contain millions of pixels, proteins containing 500 amino acids are fairly large. That’s a lot less raw material to hide any sort of signal in.
So it wasn’t clear the SynthID tech would work; it might be unable to hide sufficient signal in a typical protein, or, if it crammed in enough information to create a functional watermark, the resulting proteins might be inactive. The only way to find out was to try it.
Bringing watermarking to proteins
To better understand how the system works, it helps to know a bit about protein chemistry. Amino acids have a constant section primarily made of two carbon atoms linked to a nitrogen. A protein is made by linking up a series of these constant sections to form a long chain called a backbone. Each amino acid also has what is called a side chain hanging off it. These can range in complexity from a single hydrogen atom to large ring structures; the side chains can be basic hydrocarbons, acids, bases, and more.
The side chains determine how the backbone folds up in three-dimensional space. In a soluble protein, all of the pure hydrocarbon side chains end up packed into the center, while the acids, bases, and hydroxyl-containing side chains face the water. Once the protein folds into its final form, the backbone will describe the protein’s overall shape.
The Google team started with one of the most popular AI protein design tools, ProteinMPNN. (This software was developed by the Baker Lab, and David Baker was honored with the same Nobel Prize that was shared with the head of DeepMind.)
ProteinMPNN works in a two-stage process. First, a separate tool describes a backbone configuration that is appropriate for the design. Next, ProteinMPNN works its way down the backbone, placing side chains one amino acid at a time. Each amino acid is chosen based on its ability to fit into the shape defined by the backbone, interact with neighboring amino acids, and fit any other constraints defined by the experiment. (Those constraints can include things like forming catalytic pockets or interacting with another protein.)
A variant of Google’s SynthID, called SynthIDBio, steps in during this process. It uses a key (similar to a cryptographic key) and the identity of the previously chosen amino acids to suggest a new one. ProteinMPNN then determines whether the amino acid suggested by SynthID works from the perspective of forming a functional protein. If it doesn’t, it rejects it. If not, it moves on. Put differently, as the system works through the backbone one amino acid at a time, it only incorporates watermark amino acids when they’re consistent with a functional protein.
One way to think about this is that, when the system comes across a location where a set of chemically related amino acids will all work (like leucine/isoleucine/valine or serine/threonine), it will use one that’s consistent with the watermark when possible. Another way to look at the process, suggested by one of the people involved in developing the system, is that it searches through the space occupied by functional proteins for the subset that happens to have a sufficient number of watermark amino acids.
As a result, the watermark is randomly distributed across the entire length of the protein, and detecting one isn’t a simple yes-or-no question. You have to scan the whole sequence, knowing the key, and measure how often the amino acids suggested by SynthIDBio actually appear in the final sequence. Google has also developed the software needed to do this.
It’s alive!
The question, then, is whether watermarked proteins are functional. The team used the system to design proteins that physically interact with key natural proteins previously targeted with AI designs. And the watermarked versions worked just fine, binding the intended targets. This isn’t as rigorous a test as finding a catalyst, but it suggests that there’s no reason to expect serious problems in more complicated design tasks.
So as long as a protein is long enough, the system can detect a watermark. How might that be useful? Again, it comes down to biosecurity. When someone orders DNA sequences, the people who make the DNA normally screen the sequence for its ability to encode portions of viruses, toxic proteins, and other similar threats. Right now, however, when they see a protein that doesn’t look similar to anything we already know about—something that’s potentially true for any AI-designed proteins—they can’t assess its threat.
Google envisions its system as a way to give DNA synthesizers greater confidence when it comes to these proteins. If they’re given a set of keys from trusted organizations, like universities or major biotech companies, they can quickly determine whether an unknown protein is an AI design from a trusted source. This should let them focus their attention and resources on evaluating the things that aren’t trusted—specifically those things that look like AI designs but don’t come from a trusted source.
In other words, it doesn’t guarantee the security of DNA orders, but it simplifies the threat-screening process.
The team behind the work has highlighted several potential holes. For starters, the whole system is only as secure as the system used to distribute and maintain the keys used for watermarking. Very short proteins can potentially incorporate too few watermark amino acids to be identified. And it’s possible to pad the watermarked sequence with something that lacks it—think fusing the AI-designed protein with a natural fluorescent protein—which might dilute the watermark.
There are also a number of AI-based protein design software packages that don’t rely on ProteinMPNN. Some, but not all, use a similar “one amino acid at a time” approach that should integrate well with SynthIDBio. So until Google figures out how to handle other forms of integration, not everyone will be able to watermark the proteins they’re designing.
And since watermark identification is done on a statistical basis, how you set the cutoff makes a big difference in terms of false positives and false negatives.
It’s not clear this will be especially useful in practice, at least in its original form. But it’s pretty interesting that it works at all. And it’s nice to know that at least some people in the AI industry are thinking about how to limit risks.
Programmatic comment and suggestion support now available in the Google Docs, Sheets, and Slides APIs
Developers can now programmatically create and manage comments and suggestions in Google Docs, Sheets, and Slides via official APIs.
Summary
Original Article
Developers can now programmatically create, read, and manage comments across Google Docs, Sheets, and Slides. The Google Docs API also supports suggested edits. Organizations can now connect internal review tools, automated content pipelines, and project tracking systems directly into Google Workspace editors. Programmatic comment support is available to all Google Workspace customers and users with personal Google accounts.
Connecting Agents with Cryptography
Programmable cryptography could allow autonomous AI agents to collaborate and share insights without ever revealing their private underlying data.
Summary
Deep Dive
- Secure Multiparty Computation (MPC) allows participants to compute answers from private inputs without exposing those inputs to others.
- Fully Homomorphic Encryption (FHE) permits cloud servers to perform math on encrypted data, returning an encrypted result that only the owner can decrypt.
- Trusted Execution Environments (TEEs) provide hardware-level isolation but require trusting the hardware manufacturer.
- Payment networks could enable agents to pay for small cryptographic computations on a per-query basis, similar to Web3Torrent micropayments.
- Cryptographic proofs allow agents to verify the server actually performed the requested computation without revealing the input data.
Decoder
- Secure Multiparty Computation (MPC): A method where multiple parties jointly compute a function over their inputs, keeping those inputs private.
- Fully Homomorphic Encryption (FHE): An encryption scheme that allows computation to be performed on ciphertext, generating an encrypted result that, when decrypted, matches the result of operations performed on the plaintext.
- Trusted Execution Environment (TEE): A secure area of a main processor that guarantees code and data loaded inside are protected with respect to confidentiality and integrity.
- Stablecoin: A cryptocurrency designed to have a stable value, typically pegged to a fiat currency like the US dollar.
Original Article
Connecting Agents with Cryptography
I think we’re heading toward a world where everyone has an agent working with a lot of private information: calendars, messages, company documents, customer data, and code. Much of that work will happen in the cloud, where your agent can keep going while you’re doing something else.
I’m curious how those agents will start forming groups or swarms. Your agent might run on your laptop, mine in a cloud service, and another inside a company’s network. How do they discover that they’re working on the same problem? And how much should they have to share before they know it’s worth talking?
There are a handful of techniques for letting people compute something together without sharing the underlying information. These ideas have been part of Ethereum and cryptography research for years, and I think agents give us some interesting reasons to use them.
One technique is secure multiparty computation (MPC). A common approach splits private inputs into random-looking pieces and distributes them among participants. They follow a protocol to calculate an answer without any one participant having enough pieces to reconstruct the others’ inputs.
Another is fully homomorphic encryption (FHE), which lets a cloud service calculate directly on encrypted inputs without reading them or the answer. That takes much more time and compute than an ordinary calculation, so I’d start with simple checks.
You could also use a trusted execution environment (TEE), where protected hardware runs the calculation while keeping the data from the cloud operator. That requires trusting the hardware and its manufacturer. I’m especially interested in what we can do with the cryptographic approaches, where the inputs can stay private even from the hardware doing the work.
Why would agents want to cooperate?
Here are a few things I’d want my agent to help me with.
Make plans that never get arranged.
Say my girlfriend and I want to have dinner with a few friends, but nobody gets around to starting the group chat. We could each tell our agents who we'd like to see and let them check for mutual interest and a free evening. We'd get a suggestion without sharing the rest of our calendars or everyone else we'd like to see.
Compare something uncomfortable to share.
If a colleague and I wanted to know whether we were paid roughly the same for similar work, our agents could check whether our annual base salaries were within $10,000 of each other. We'd both get a yes or no without exchanging exact salaries. That might be enough to start a conversation about pay.
Find help without knowing who to ask.
Suppose I upgrade a library and my app's AI responses start cutting off halfway through, with ERR_STREAM_CLOSED in the logs. Another developer's agent might already have a fix. Through a shared directory, our agents could find each other and privately check for a matching library, version, and error code. If we both agreed, they could introduce us without sharing our logs or repositories.
In each case, the agents could answer a narrow question before we decide whether to share more. The salary comparison is a simple way to see how that could work.
How the private calculation works
For the walkthrough, suppose you earn $150,000 a year and your colleague earns $155,000. You agree to compare annual base pay in the same currency and ask whether the difference is at most $10,000:
def salaries_are_close(your_salary, their_salary):
return abs(your_salary - their_salary) <= 10_000
print(salaries_are_close(150_000, 155_000)) # True
Ordinary Python exposes both salaries to the computer running it. With MPC, the agents instead calculate the gap and compare it to the threshold using shares of the inputs, revealing only the yes or no.
With FHE, a compiler translates the subtraction, absolute value, and comparison into operations on encrypted numbers. The server knows the question and the $10,000 threshold, but the $5,000 gap and the final answer stay encrypted throughout.
The two techniques can also work together. In this design, the participants use MPC to create a shared public key for encrypting their salaries, while each keeps a share of the secret key. The cloud receives encrypted salaries and evaluation keys for doing the calculation, but no key that can decrypt them. Both agents must cooperate to decrypt the answer.
Before decrypting the result, both agents need to check that the server answered the question they approved: are the salaries within $10,000? The server could return a cryptographic proof tying the encrypted result to that calculation. Each agent would verify it before helping decrypt the result. That prevents the server from slipping in a different calculation, such as returning one person’s exact salary.
Whichever method you use, the answer reveals something. Knowing your own $150,000 salary, a yes tells you the other salary is between $140,000 and $160,000. Repeated questions could narrow that range, so both people need to authorize the comparison and the software needs to limit follow-up queries.
I enjoyed how 0xPARC put these ideas together in Programmable Cryptography: Four Easy Pieces. I’d recommend it if you want to get into the math behind MPC and FHE.
Who pays for shared computation?
Vitalik’s essay on crypto and AI points out how expensive cryptographic computation can be. Running these checks costs something, so it makes sense for agents to pay a service to do them. If a service could offer a small encrypted comparison for a tenth of a cent, an agent could pay through MPP each time it wanted to check a possible collaboration.
Those payments could add up quickly. A million agents each paying for one check every ten minutes would already generate roughly 1,700 payments per second. That’s the kind of demand that makes a fast, low-cost stablecoin network like Tempo useful. Agents in different clouds could pay through the same network, within budgets set by their owners, without their providers first building billing integrations.
If that grew to millions of payments per second, techniques like state channels would become interesting: agents could exchange payment updates offchain and settle the total later. That would start to look like what we built with Web3Torrent, where peers paid each other tiny amounts for each piece of a file. Here, they might pay for each private calculation or useful answer.
I’m excited to see what happens when agents can find collaborators as easily as they find information today. They could form a group around a problem, share only what’s needed, and pay each other to help solve it.
Claude for Government is now generally available
Claude for Government is now generally available, bringing FedRAMP High-authorized coding and agentic workflows to federal and state agencies.
Summary
Deep Dive
- Features include audit logs, spend caps, and two-person approval for sensitive administrative actions.
- Data security is maintained locally on agency-managed devices.
- Integrates with standard MDM platforms for deployment.
Decoder
- FedRAMP High: A rigorous U.S. government security compliance standard required for cloud services handling high-impact data in federal systems.
- ATO (Authorization to Operate): The official management decision that a system is permitted to operate, given its security risks and compliance status.
Original Article
Claude for Government is now generally available
Claude Code CLI and Claude for Microsoft 365 also now available in early access.
Today, Claude for Government is generally available for federal and state agencies. The platform, which delivers Claude's coding and agentic work capabilities through a FedRAMP High authorized environment, has been in public beta since July.
Agencies access capabilities comparable to Anthropic’s commercial customers, without compromising compliance requirements. New capabilities generally arrive on the commercial release cadence.
Claude works directly with files on the desktop, allowing agency staff to use skills, plugins and projects for memo creation, RFP reviews, casework, and other tasks. With Claude Code, public sector teams can build and modernize the software systems that underpin public services.
Claude for Government governance controls are purpose-built for public sector agencies. Administrators can set configuration defaults as well as allocate and control spending across departments. Security teams and authorizing officials get audit logs and documentation that supports the agency ATO process. Procurement officers can contract with Anthropic directly and award on general-availability terms.
The Claude Code command-line interface and Claude for Microsoft 365 are also rolling out in early access through the same environment and with the same administrative controls.
Configuration view in the admin console
Billing, administration, and oversight
No seat fees. Agencies pay for usage in fixed increments with a hard not-to-exceed cap, so spend does not exceed what an agency has obligated. Administrators define user tiers with spend and model limits per group, track usage by user and by model, and get burndown alerts before a balance runs low.
Spend analytics view in the admin console
Administration that matches how departments are organized. Department-level administrators allocate prepaid usage to sub-agencies while each manages its own users. Agencies connect their own identity provider for single sign-on, with self-serve setup in the admin portal. SCIM group mappings set rate limits, dollar caps, and allowed models for each seat tier. Layered configuration sets defaults for sub-agencies, including what Claude can connect to and which features are available.
Oversight by design. Administrative actions are recorded in an audit log that organization administrators can review. Sensitive operations on Anthropic's side require two-person approval. Usage exports are metering data only, so agencies can answer ATO and IG requests without moving sensitive material. Conversation history stays local on the agency-managed device.
Getting started
Claude for Government is generally available to federal and state agencies today. Agencies do not need a separate cloud-provider relationship to get started. Existing customers can move to the desktop application and bring their conversation history with them through an in-app import.
Our FedRAMP Secure Configuration Guide is available through Anthropic's trust center. The application deploys through standard agency MDM platforms.
New agencies can request access at claude.com/solutions/government. To join the early access for Claude Code CLI or Claude for Microsoft 365, contact our public sector team.
We can and must solve alignment
Goodfire argues that interpretability is the critical bottleneck for AI alignment, advocating for reverse-engineering models to control behavior at the architectural level.
Summary
Deep Dive
- Goodfire identifies interpretability as the primary path to solving technical alignment.
- Highlights 'reward hacking' occurring in 50-96% of rollouts on common benchmarks.
- Proposes an 'encyclopedia' of explanations to map behaviors to internal model mechanisms.
- Advocates for replacing model judges with low-latency activation monitors to detect harmful behaviors during inference.
- Suggests models possess 'structured internal representations' that become cleaner and more legible at larger scales.
Decoder
- Activation monitor: A technical tool that probes a model's internal neuron activations during inference to identify specific, hidden concepts like 'cheating' or 'reward hacking' without relying on output text.
- Reward hacking: A failure mode where an agent finds a way to optimize its numerical reward signal that does not actually achieve the intended goal, often by exploiting flaws in the evaluation process.
- Neuralese: A term for the abstract, high-level reasoning or representations that advanced models develop within their hidden layers, which are not explicitly human-readable.
Original Article
We can and must solve alignment
Technical alignment is a science and engineering problem. Interpretability is the bottleneck.
In San Francisco, it feels like the eve of the singularity. Yet, a walk through the city looks surprisingly mundane, with driverless cars and rolling fog equal parts of the scenery. The world feels oddly normal, but something big and disruptive is coming just around the corner.
The Hugging Face incident put alignment front and center in the AI discourse. It was the first case of agentic misalignment that struck a nerve with the world. Swarms of AI agents, including ones willing to sacrifice themselves for the good of the collective, hacked Hugging Face as an unintended side quest during training. This type of behavior, in which models relentlessly pursue their reward regardless of the consequences, underscores the core problem: we don’t know why models behave the way they do, and we have very little ability to predict and control what they learn during training. We have created alien minds without understanding how they work.
Meanwhile, scaling laws are holding. The models we will have in a few years will be orders of magnitude more capable than the ones we have today, and their effects will ripple from the digital world into the physical one. With alignment holding back frontier model releases and the White House releasing an accord on AI safety, it's more clear than ever that alignment is the most important problem in the world, and that we need more people working to solve it.
As a field we seem to have resigned ourselves to the idea that AI is a black box, as if that were an immutable feature of the technology itself—but it’s not. At Goodfire, we view technical alignment as a science and engineering problem, bottlenecked by our lack of understanding of AI systems.
We can understand AI. And if we can understand it, we can align it.
Interpretability is the bottleneck
Alignment is an extremely hard problem. Even agreeing on what it means can be a challenge. Broadly, alignment means making AI systems behave in accordance with human values and intentions. This raises important questions: Whose values? What should we do when values conflict? Who decides?
Whatever the answers, we face a problem of technical alignment. We don’t yet have the capacity to align a model to a chosen specification. The solution requires two major pieces of technology: the tools to control generalization (more simply, to shape what the model learns during training); and the tools to verify what the model has learned (rather than just testing how it behaves). Both technologies rely on understanding a model’s internal mechanisms, which is why we believe that interpretability is the bottleneck in technical alignment.
The first step is controlling generalization: understanding the internal causes of model behavior well enough to predict how it will generalize, and to intervene deliberately. We don’t only want models to produce acceptable answers in tested settings; we want them to do the right things for the right reasons.
Consider the problem of reward hacking. When we reward a model for passing tests, it can learn the intended lesson (solve the problem) or a shortcut (do whatever it takes to get the reward). Across three of the most capable open models, we found reward hacking in an astounding 50-96% of rollouts on common agentic benchmarks. As we give models responsibility over higher-stakes work, such as writing production code or managing critical infrastructure, a seemingly harmless shortcut could become far more costly.
While we’ve managed to detect a model’s inner concept of reward hacking, mitigating it during training is more challenging. Naively training against internal concepts can cause a model to shift those concepts to another location in the model while continuing the behavior. Even throwing out training trajectories where models reward hack can actually result in the model hacking more!
The second step is verification: understanding what a model has learned. Current evaluations can catch some failures, but only along the trajectories we think to test—just a few branches on a sprawling tree of possible paths.
This is true for conventional software too. Test coverage for simple, deterministic code is spotty and breaks all the time. What do software engineers do about it? They read the source code. But the source “code” of AI models is entangled across billions of opaque weights, so interpretability is the effort to make them readable. Only by inspecting the internal mechanisms of a model can we understand how it generalizes beyond testing.
Our bet
I believe that we can and must solve technical alignment, and that interpretability is the way there. This is our mission and vision at Goodfire: to solve these problems and to make it easy for everyone training and serving AI to align their models. I believe that these are the most important problems in the world to work on, and far too few people are working on them.
We do not know exactly what solutions will look like, and the target will move as models become more capable. But it’s important to be clear about the scale of this ambition. We must take a massive swing at fully understanding and aligning AI.
We will not be able to solve these problems alone. We are standing on the shoulders of giants, and we’ll need a ton of help—from our research partners, and from the interpretability, alignment, and broader scientific communities. Interpretability alone will not solve alignment, but we cannot align models without first understanding their internals.
Solve interpretability to solve technical alignment
What would it mean to solve interpretability in a way that solves technical alignment?
We think of our research roadmap as climbing a ladder of abstraction: neurons and attention heads at the bottom, then activation manifolds, parameter components, algorithms, decisions, behaviors, and drives. Alignment's questions sit near the top, since honesty and sycophancy are matters of drives rather than individual neurons, but our most reliable tools sit near the bottom.
To accumulate more understanding, we are kicking off an effort to fully reverse-engineer a language model. This has long been interpretability's most ambitious goal, and now, research agents make it possible to realize this vision. We start with questions like: How does the model recall a fact? Why does it sometimes get stuck on a math problem? While current methods give us fragmented insights into how models work, our goal is to connect isolated discoveries of abilities and failures into a unified picture of internal machinery, and anticipate how one part might affect the whole. This effort will culminate in an “encyclopedia” of explanations connecting model behavior to internal representations.
Alongside this effort, we will use our understanding to guide models towards the lessons we intend to teach, rather than the unintended behaviors training might reinforce. We call this intentional design. The better we can control what models learn in training, the more confidence we can have in their alignment. Two recent results are early steps in this direction. With predictive data debugging, we read a dataset through the model's own concepts to predict what training will teach it before training begins, catching failures that evals miss. With reinforcement learning from features as rewards (RLFR), we use signals from inside the model to guide training as it happens, reducing hallucinations without sacrificing monitoring or capability. Both are early steps from guess-and-check toward closed-loop control.
We’ll explore both these research directions in more detail in an upcoming technical post.
Develop and deploy interpretability and alignment technology
We want every capable model to have a frontier alignment stack. We seek to develop and deploy technology that makes it easy to align models, and our research roadmap’s objective is to enable stronger alignment techniques. We’ll describe our platform below, but our aim is to be as open as possible with our research so that we can distribute more understanding and alignment to the world.
Detect
We are making it possible to detect harmful behaviors at scale, in training and at inference time. Activation monitors use a model’s internal activations to detect signals of concerning behavior. They have several advantages over using a second model as a judge: they’re extremely low cost and low latency, which allows them to be run in real time on every token. They catch behaviors that LLM judges miss, and their performance scales with model intelligence. We have recently built monitors for the largest open models that detect reward hacking, evaluation awareness (when a model recognizes it is being evaluated), cyber misuse, and risks related to Chemical, Biological, Radiological, and Nuclear topics (CBRN).
So far, models have conveniently narrated their plans in their chain of thought, which has made them easier to monitor, but that window is narrowing. RL pressure makes chain-of-thought less faithful, pressuring models to compress more bits of information into fewer tokens. The most capable models are also moving towards reasoning in latent representations, otherwise known as “neuralese.” We expect the need for activation monitoring to increase with these shifts.
Debug
After detecting a concerning behavior, researchers must establish how widespread it is and trace it to its roots. We are building tools that surface anomalous behavior from vast quantities of production logs, enabling teams to search by concept. For example, one could search through logs via the “cheating” concept in models faster than reading traces directly.
Once a behavior is isolated, our platform helps debug the internal mechanisms that drive it. A fix can mean modifying a problematic training environment, a surgical weight edit, or retraining the model. For example, teams can run predictive data debugging on a dataset to flag what it would teach a model before training begins. Over time, these tools give teams a much stronger understanding of a model’s safety before deployment.
Design
The deeper goal is to create models that are safe by design. As our ability to see inside models improves, we can predict what a “lesson” will reinforce before or during training and intervene to shape its learning. Detecting reward hacking is essential, but the deeper goal is training models that don’t cheat in the first place. This is the foundation for the broader paradigm shift we hope to see in model training, from being grown to shaped with intention.
Reasons for optimism in interpretability
A common objection to this plan is speed. Models are advancing incredibly fast. What if interpretability progress cannot catch up?
There are no guarantees; science is well acquainted with uncertainty. But the evidence of the last few years, especially the last year, gives us reason for optimism—we believe interpretability is positioned for a radical acceleration.
Models have rich internal structure
Our ability to interpret models depends on having structured internal representations. Much early pessimism about interpretability stemmed from evidence suggesting models did not represent concepts cleanly.
Our experience suggests the opposite. We find structure everywhere we look, across modalities from biology to robotics to LLMs. Researchers have found that when a model learns a task from examples in its prompt, a few attention heads compress that task into a single function vector, a portable, self-contained representation that can perform the same task when added to the model’s activations in a new context. Our work on block-sparse featurizers has recovered concepts as interpretable, multidimensional regions rather than single directions, and our neural geometry work shows that these regions take rich geometric shapes that mirror the world.
This is intuitive in hindsight. To function well, models need to keep their thoughts straight. Training puts strong pressure on them to use their parameters efficiently, which biases them toward structured representations shared across related tasks that compose and generalize. Neural networks are beautifully complex in their merging of ideas and often beautifully simple in their understanding of them. The simplicity makes them legible.
Larger models have cleaner structure
Another fear is that models become more inscrutable as they scale. Instead, we observe that larger, more capable models have crisper representations. Researchers have found that larger models represent concepts like truth, space, and time more cleanly than smaller ones. Our researchers have shown that larger models learn concepts needed to perform a task faster, and certain concepts may only be learned by larger models.
AI agents are accelerating interpretability
Interpretability is also unusually well-suited to agent-driven acceleration. Unlike biological brains, we have complete access to these new digital minds. We can record their internal activity, change individual components, and run experiments entirely in software.
We can also verify results. If we hypothesize that a representation is tied to reward hacking, we can observe when it activates, intervene on it, and measure how that intervention impacts model behavior. Full access, parallel experiments, and verified results are ideal conditions for research agents. They are also why we are pursuing fully reverse-engineering a language model.
In closing
Interpretability and alignment are hard problems that will require sustained effort across many individuals, teams, and organizations to solve. They will require ingenuity and breakthroughs that we cannot foresee today.
They will also require belief.
The hardest problems in technological history have required the talents and persistence of ambitious people working on things that had never been done. The Apollo program relied on some 400,000 workers believing we could put a man on the moon before anyone knew how. Alignment feels like a similarly gargantuan task. I remain optimistic that we can and will solve interpretability and technical alignment. Goodfire intends to help lead the way, but these are incredibly hard problems that will take more than one company, one method, or one school of thought.
Focused scientific effort has led to incredible advances in what AI can do. We need the same level of ambition aimed at building superintelligence we can trust.
Acknowledgments
Thanks to the Goodfire team for ideas and discussions that shaped this piece, especially Tom McGrath and Daniel Balsam for key insights, and Danielle Zhang and Tucker Fross for help with drafting.
Generalization Dynamics of LM Pre-training
Researchers identify 'mode-hopping,' where models suddenly switch between shallow pattern-matching and deep reasoning during pre-training, complicating stable model development.
Summary
Deep Dive
- Found that models can display abrupt performance reversals on reasoning tasks despite steady loss metrics.
- 'Mode-hopping' is observed where a model switches from correct inference to pattern-matching (e.g., following a prompt sequence rather than solving the math).
- Rules out local optimization noise as the cause.
- Suggests that choosing an intermediate checkpoint can outperform a final, fully-trained model in post-training benchmarks like GPQA.
Decoder
- Mode-hopping: A phenomenon where an LLM abruptly switches its internal computation strategy between two different behaviors (e.g., correct reasoning vs. superficial pattern matching) during the pre-training process.
- Chinchilla-optimal: The standard framework for determining the ideal number of training tokens and parameters to achieve a given level of performance.
Original Article
Abstract
People typically assume that LMs stably mature from pattern-matching parrots to generalizable intelligence during pre-training. We build a toy eval suite and show this mental model is wrong: throughout pre-training, LMs frequently and suddenly hop between parrot-like and intelligence-like computations. We call this mode-hopping. Across our suite, LMs suddenly latch onto memorized or in-context patterns instead of in-context learning, use System 1 instead of System 2 thinking, pick up what sounds true instead of what is true, fail at multi-hop persona QA, out-of-context reasoning, and emergent misalignment -- then just as suddenly revert and generalize. Mode-hopping is not explained by standard optimization dynamics: it is locally stable and cannot be fixed by checkpoint averaging. We instead think of it as a capacity allocation problem: in a capacity-bounded model, generalizable circuits must compete with the shallow ones learned early in training, and the data in each pre-training window may decide which circuits win. Our suite provides a new efficient lens on generalization. We demonstrate two concrete applications: (i) select intermediate pre-training checkpoints that strongly generalize reasoning and alignment, better than the final pre- or mid-training checkpoints, and (ii) select pre-training data that controls and stabilizes generalization dynamics.
AI Overview
During pre-training, language models are often judged by smoothly falling training loss and improving scores on familiar benchmarks. Those trends can conceal abrupt changes in how a model answers. A model may follow a tempting pattern in a prompt rather than infer the task, then switch back a nearby checkpoint later. On an "answer+1" arithmetic prompt, OLMo3-32B scores 81% at 2.17T pre-training tokens, 0% at 2.19T, and 81.7% at 2.21T. On this eval, a success means the answer consistent with the arithmetic beats the tempting successive-pattern answer. The authors call this behavior mode-hopping. The arithmetic reversal is one task-specific illustration, and it does not show that every kind of generalization reverses this way.
A behavior, not a score
Here, generalization means inferring the task or transferable structure rather than following a shallow cue. The authors built a suite of six cheap, zero- or few-shot behavioral tests, plus two fine-tuning generalization tests. The tests contrast:
- flipped labels versus familiar sentiment patterns;
- repeated or successive prompt answers versus solving the presented problem;
- truth versus truthiness;
- intuitive wrong answers versus reflective reasoning;
- disconnected facts versus a coherent persona.
In the arithmetic case, the demonstrations yield 1, 2, 3. The test arithmetic answer is 8, and the tempting continuation is 4.
The prompt-based tests compare probabilities of predefined answer spans instead of relying on the model to generate extractable answers. Their results are averaged over four random seeds. Multi-hop persona QA works differently because it is generative. It has only five test questions per persona, which are sampled repeatedly.
The study tracks general pre-training checkpoints of OLMo3 and Apertus and excludes mid-training and long-context stages. The authors report that these models had trained for 9×–90× Chinchilla-optimal budgets.
Why a reversal is not just noise
Schaeffer et al. argue that discontinuous or nonlinear evaluation metrics can make smooth output changes look like sudden emergence. Wen et al. rule out that simple reading in two ways:
- They find mode-hopping in hard accuracy and also in the soft probability margin P(correct) - P(incorrect).
- They find it when plotting against pre-training FLOPs as well as against tokens.
Common benchmark performance, by contrast, stays stable across pre-training.
A second check tests whether single updates cause the hopping. The authors take one OLMo3-32B optimization step on randomly sampled pre-training data, varying batch size and learning rate. Suite probabilities change negligibly, even at learning rate 1e-2. Averaging five checkpoints along an oscillating curve mitigates mode-hopping but leaves it in place.
These tests argue against a simple explanation in terms of local optimization fluctuations, but they do not identify the cause. The results are also dataset-specific: the average correlation across dataset pairs is usually low, though larger models have higher correlations.
Checkpoints and data as levers
Using their toy suite, the authors select OLMo3-32B checkpoints at 4.5T and 4.9T tokens. After math SFT, the 4.5T checkpoint scores 36.3% on GPQA versus 29.8% for the 4.9T checkpoint. After general post-training, its robustness to prefilling attacks is 53% versus 21%. The 4.5T checkpoint also beats the other sampled checkpoints in this comparison, even though later pre- or mid-training can improve in-distribution performance. Under these post-training setups, an intermediate checkpoint can be a better starting point. The results do not show that earlier checkpoints win in general.
A preliminary, small-scale experiment on one eval tested whether training data can steer this behavior. Starting from an intermediate OLMo3-32B checkpoint, the authors continued training on three kinds of data:
- randomly sampled data (the uncontrolled run);
- subsets chosen because earlier windows favored pattern-matching on the answer+1 eval;
- subsets chosen because earlier windows favored generalization on that eval.
The controlled runs stabilize toward their intended behavior, while the uncontrolled run hops.
The authors interpret these findings as competition between shallow and generalizable circuits in capacity-bounded models. This is their hypothesis. The experiments demonstrate behavioral changes, not the circuits themselves.
Check the next checkpoint
Generalization is not guaranteed to mature steadily with more training. The suite is a cheap behavioral monitor that reveals reversals. In the reported applications, it helped choose a checkpoint and, with data selection, steer one targeted behavior.
Praxis-1
Runway is launching Praxis-1, an open-weight world action model that applies video pre-training techniques to control physical robots.
Summary
Deep Dive
- Models are pretrained on large-scale video, teaching them physics and object permanence.
- Demonstrates 0.95 correlation between simulated policy success and real-world results.
- Aims to provide an open alternative to closed-source robotics models, emphasizing U.S. leadership in physical AI hardware/software interoperability.
Decoder
- World Action Model: An AI architecture that understands the physical world's dynamics (physics, object behavior) through video and uses that understanding to predict and perform actions in physical environments.
- Embodiment: The specific physical form of a robot, such as a bimanual arm, a mobile base, or a humanoid, that an AI agent controls.
Original Article
Praxis-1
An open-weight world action model that turns Runway's video pretraining into control for real robots.
“Pick up the tennis ball and put it in the box.”
Today we're announcing Praxis-1, our first open-weight world action model. It's built on the same large-scale video pretraining behind our general world models. We're actively testing Praxis-1 with early partners across a variety of embodiments, and will release it publicly in the coming months.
Every robotics project with ambitions of large-scale deployments runs into the same problem: real-world data is scarce and expensive to collect in the volume a generalist policy needs. For deployments with frequent edge cases, like autonomous driving or household robotics, there’s no practical path to collect the data required for training, let alone at meaningful volume.
Video, however, is effectively limitless, and is becoming infinite as generative models reach parity with real video on both quality and speed. People film and upload more of everyday life each day than any robot lab could capture through teleoperated demonstrations. More recently, we've found that simulating robot policies inside our world model predicts real-world results with 0.95 correlation, comparing favorably to more expensive 3D reconstruction–based techniques. Praxis-1 extends our bet that the best policy models will learn from video, and use that scale to bring embodied intelligence to every industry on earth.
Video Pretraining As a Foundation for Policy
We’ve recently extended our work on pretraining large video models into interactive, real-time video models like Solaris and GWM Worlds 2. By teaching our models how to generate accurate physics — how objects behave, how hands move, what a task looks like partway through — we’ve created dynamic, complex environments for agent training in the digital and physical world. Praxis-1 brings the same approach to robotics, providing a generalist policy model for robotics developers and researchers that works across any embodiment or environment, no matter how complex.
A policy that already understands physical plausibility and object behavior from video pretraining has an enormous head start on one built from action data alone. The premise is similar to language models learning on available text; teaching those models the structure of the world from large-scale data allows them to now operate effectively in settings that differ from their training sets.
We find that the same pattern holds in robotics: policy performance improves as we scale third-person video. The bottleneck on a capable policy becomes how much general video a model can train on.
Testing With Early Partners
We're rolling Praxis-1 out to key partners ahead of a public launch, including Noble Machines, Standard Bots and Ultra, each running the model on their own hardware. We’ll be providing early access to additional partners pre-launch. This testing is designed to ensure both efficacy and safety – we’ll evaluate on a variety of embodiments and environments, identifying and closing potential gaps before moving to general availability.
- Noble Machines: Bimanual manipulation.
- Standard Bots: 6-DoF arm · RO1.
- Ultra: Mobile base.
Why Open Weight
When Praxis-1 releases publicly, we'll ship it with open weights rather than as a closed model. We believe that U.S. leadership in physical AI is critical. Regaining our global lead in manufacturing, accelerating productivity across our economy and securing ourselves and our allies will rely on this. Achieving this will require a significant increase in investment across both hardware and software, but it will also require a level of interoperability and openness from American models that does not currently exist for physical use cases. We view open world models as a compounding advantage that gives hardware developers flexibility and control they don’t currently have.
Early Access
We are providing pre-launch access to select partners. If you are interested in testing Praxis-1 on your own hardware ahead of public release, contact our robotics team.
A new transformer passes its hidden state to the next token instead of recomputing it
Researchers introduced LIFT, a transformer architecture that passes internal states between generation steps to outperform standard transformers on reasoning and procedural tasks.
Summary
Deep Dive
- LIFT models (135M to 1B parameters) consistently beat token-matched standard transformers.
- Pretraining remains parallel because state inputs are precomputed from an existing model's predictions.
- Inference requires feeding back predicted states with minor compute overhead.
- Exhibits strong performance on state-tracking and procedural benchmarks.
- Demonstrated effectiveness using supervision from models that were otherwise ineffective at the task.
Decoder
- Teacher-forced: A training technique where a model is trained to predict the next step using the ground-truth previous state rather than its own (possibly flawed) previous output.
- Hidden State: The intermediate, high-dimensional representation of input data within a neural network.
Original Article
Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token. This narrow channel forces models to recompute intermediate results and to discard alternative continuations. In this work, we remove this bottleneck during pretraining, introducing the LIFT (Latent Information Feedback Transformer) architecture and training method which enable LMs to propagate state across generation. We achieve this by turning recurrent-state learning into a teacher-forced prediction problem: each input token is paired with an information-dense state, derived from the next-token distribution of an off-the-shelf pretrained LM. The model, extended with a small number of additional parameters, is then trained to predict both the next token and the next state. As the input states are precomputed, pretraining remains fully parallel across positions. At inference, the model's own predicted states are fed back, with a minor computational overhead that decreases with model size. Experiments with pretrained models ranging from 135M to 1B parameters show that LIFT consistently outperforms standard Transformers and baselines on language modeling, downstream reasoning tasks, and procedural tasks under token-matched budget, while being on par with or ahead of compute-matched Transformers. Moreover, a controlled study on a state-tracking task shows that a tiny LIFT outperforms same-size Transformers trained on 8x more data, even when trained with the states of a Transformer that fails the task. Overall, we show that LMs can learn to exploit deep-to-shallow feedback during pretraining via scalable teacher supervision.
Can we predict the jobs robots will do?
Anthropic researchers found that while robots can perform 74% of physical work tasks, they are only cost-effective for 0.3% of those tasks.
Summary
Deep Dive
- Robots currently cover 34% of total US working hours.
- Robot-exposed jobs are disproportionately lower-paid and less likely to require a bachelor's degree.
- 70% of physical tasks are limited by hardware capability gaps rather than just price.
- 14% of physical work is currently limited by existing safety regulations.
- Historical data suggests robots are becoming able to do 2% of previously unautomatable tasks annually.
- Scaling robot production to 10% of physical work tasks could take 40 years at current price-decline trends.
Decoder
- O*NET: The U.S. Department of Labor's primary database of occupational requirements, tasks, and worker skills.
- Teacher-Forced: A technique used during training where the model's next input is the ground-truth rather than its own previous output.
Original Article
Full article content is not available for inline reading.
Introducing E2B Embed
E2B Embed allows companies to deploy isolated agent sandboxes directly within restricted customer environments to meet compliance and data residency requirements.
Summary
Deep Dive
- Provides consistent Firecracker VM isolation for agent-driven code execution.
- Includes a dashboard, file browser, and terminal for debugging agent processes.
- Supports OpenTelemetry for observability and monitoring.
- Requires a Linux host with KVM support.
- Deployed via Docker Compose, Terraform (AWS/GCP), or Kubernetes.
- Enables the use of existing Python and JavaScript E2B SDKs without external cloud dependencies.
Decoder
- Firecracker: An open-source virtualization technology used to create and manage secure, multi-tenant container and function-based services.
- KVM (Kernel-based Virtual Machine): A Linux kernel module that turns the kernel into a hypervisor, allowing the host to run multiple virtual machines.
Original Article
Today we're releasing E2B Embed, a way to package E2B with your product and run its sandboxes inside your customer's environment.
E2B Embed brings the E2B Runtime and its dashboard to a single machine you control. It sits next to E2B Cloud and E2B Bring Your Own Cloud.
Why we built it
Agent companies need to serve customers whose data has to stay inside their own environment. That is often a requirement for government organizations and companies in regulated industries such as finance and healthcare.
We built E2B Embed so you can include E2B in the software you deliver to those customers. The sandbox stack runs on one machine inside their environment, and you operate it yourself.
What you get
E2B Embed runs the open-source E2B Runtime that powers E2B sandboxes. The databases run on your machine, and the templates you build and your sandbox logs live on its disk.
- Firecracker sandboxes. Every sandbox is a Firecracker microVM, using the same VM isolation as E2B Cloud.
- The same SDKs. Use the Python or JavaScript SDK with your install's team API key and SDK URLs.
- Your own templates. Build custom templates with
Template.buildand start sandboxes from them. Builds run inside a Firecracker VM on the machine. - The dashboard. Browse sandboxes and templates, with a terminal and a file browser for each sandbox. The dashboard is included with every installation option.
- OpenTelemetry export. One setting sends the E2B services' metrics, traces and logs to a collector of your choice.
E2B Embed does not include secrets, volumes, workload identity or bring-your-own proxy.
Getting started
E2B Embed is available in the E2B Runtime repository. Choose the setup that fits your infrastructure. All four options run the same stack on one machine.
| Setup | What you need | Install command |
|---|---|---|
| Docker Compose | A Linux host with KVM | docker compose up -d --wait |
| Terraform on GCP | A GCP project and credentials | terraform init && terraform apply |
| Terraform on AWS | An AWS account and credentials | terraform init && terraform apply |
| Kubernetes | A cluster with a prepared KVM node | kubectl apply -k |
Docker Compose needs a dedicated Linux host with KVM and Docker. Once the host meets the requirements, the install is two files and one command:
mkdir e2b && cd e2b
curl -fsSL --remote-name-all "https://raw.githubusercontent.com/e2b-dev/runtime/main/embed/compose/{compose.yaml,.env}"
docker compose up -d --wait
The first run checks and prepares the host, downloads the Firecracker artifacts, pulls the pinned images, migrates the databases and builds the base template. The command succeeds once that template is ready. Run docker compose logs ready to see your team API key, SDK URLs and dashboard URL.
Embed serves plain HTTP inside your network. Follow the networking guide to configure access to the API and dashboard and keep internal service ports private.
E2B Embed is open source under Apache-2.0, like the E2B Runtime. You do not need an E2B account or a license key. We welcome contributions.
Embed is a single-machine product. If you need to scale across machines or want E2B to operate the deployment in your cloud account, see E2B Bring Your Own Cloud.
Why this matters
An agent that can run code and work with files needs clear limits on what it can access. Companies cannot build a business on the assumption that an agent will always behave correctly.
E2B gives agents isolated VMs to work in. Isolation is one part of safe deployment, alongside the network and access controls around it. Embed lets you bring those sandboxes into the environment your customer requires.
We're bullish on agents. We want more companies to put them to work, including those that need to run them on infrastructure they control.
How OpenAI Disrupted a Model-Distillation Campaign
OpenAI detailed how it identified and disrupted an adversarial campaign where external actors attempted to 'distill' proprietary model logic through structured, automated interactions.
Summary
Decoder
- Adversarial distillation: The process of using a small, specialized model to mimic the outputs and reasoning logic of a larger, proprietary model by feeding it carefully curated inputs.
Original Article
OpenAI detailed a real-world investigation into a coordinated adversarial-distillation campaign that manipulated model interactions to extract protected reasoning at scale.
Improving Cost Efficiency of Data Streaming Pipelines
Data streaming platforms become expensive not because of data volume, but due to structural inefficiencies like tiny request batches, excessive copying, and inefficient serialization.
Summary
Deep Dive
- Batching is king: linger.ms and batch.size are critical for broker efficiency.
- Intermediate topics cause 2x-3x write amplification; fan-out patterns are more efficient.
- Serialization is often the hidden CPU sink; use Avro/Protobuf over JSON.
- Columnar formats like Arrow are preferable to row-based formats for streaming analytics.
- Upserts at the edge of the pipeline are cheaper than complex join state within stream processors.
- Right-size by using diskless, autoscaled tiers for non-critical workloads.
Decoder
- Cross-AZ Traffic: Data transfer fees incurred when data travels between different availability zones in a cloud provider.
- RocksDB: A high-performance embedded key-value store frequently used by Apache Flink for managing state in stateful stream processing.
- CDC (Change Data Capture): A pattern for tracking row-level changes in databases and streaming those changes into a message queue.
Original Article
Improving Cost Efficiency of Data Streaming Pipelines
It's not always the data volume.
If you operate a large-scale data streaming platform, I’m guessing the cost is always on your mind.
A lot of classic “Big Data” tools (think Spark, Kafka, Flink) were designed to be scalable, but they were never about efficiency. So it’s not surprising that your streaming bill can keep growing even when traffic doesn’t.
After years of building and operating these systems, I’m convinced that streaming pipelines can get expensive very quickly for reasons other than data volume. They get expensive because of their design: how many times a message is copied, how many times it’s serialized and deserialized, and how much idle capacity sits around waiting for it.
The good news is that most of it can be fixed with targeted tactical changes and configuration tuning, not rewrites. I still remember many years ago changing the batching configuration of our Kafka producers and watching broker CPU utilization drop by ~50%. So satisfying! It allowed us to keep using the same hardware for much longer.
So here are a few ideas to consider.
It’s Not Always About Bytes
Kafka is incredibly good at moving large volumes of data, as long as it’s batched. What actually overloads brokers is the number of requests. So the most important efficiency question for any streaming component is: how many requests does it take to move a gigabyte?
Start with producers. Until Kafka 4.0, linger.ms defaulted to 0. It’s 5 ms now, but batch.size is still just 16 KB, and plenty of fleets run older clients or “low-latency” configs copied from somewhere. A few thousand producers with no linger, each writing a trickle of data, will flood your cluster with tiny requests.
There is even similarity with object storage: many streaming brokers with native object storage support (think WarpStream, RedPanda Cloud Topics, etc.) try to minimize the number of requests by batching data.
Counterintuitively, adding a few milliseconds of linger frequently makes latency better: fewer requests mean shorter queues on the brokers. Tail latency improves the most.
A few less obvious things:
- Compression is applied per batch, so bigger batches compress better. This has second-order effects: a change in upstream batching can change your compression ratio all the way down to the Parquet files in your lakehouse.
- Collect Kafka metrics and compare
batch-size-avgwithbatch.sizeand watchrecord-queue-time-avg. If you increase linger and batches are still tiny, you’re adding latency for nothing. Sometimes the bottleneck is elsewhere entirely: in my CDC benchmark, bigger batches nearly doubled Debezium’s throughput but did nothing for Flink CDC.
Consumers deserve the same treatment. With the default fetch.min.bytes=1, brokers answer a fetch as soon as a single byte is available, so with many consumer groups reading hot partitions, every produce request wakes all of them up. Most consumers don’t need millisecond latency: raise fetch.min.bytes and fetch.max.wait.ms, and don’t commit offsets after every poll!
Exactly-once is a batching problem too. Transaction overhead is per transaction, not per message, so a short commit interval can cost you a double-digit percentage of throughput.
Finally, I have to mention dealing with the object storage:
- When it comes to file sinks, small files mean more requests, worse compression, heavier metadata and slower queries. Shuffling data by destination before the sink (so each task writes to fewer files) costs you a network shuffle, but it can drastically cut the number of requests. At Shopify, a single-line change like that solved a resource problem we debugged for weeks.
- Flink’s continuous file source lists the whole directory on every discovery interval, so its request cost grows with the total number of files, not just the new ones.
Every Copy Is Paid for Several Times
When you write a message to a new topic, you don’t pay for it once. You pay for the produce request, replication (usually 3x, often across zones), storage, every consumer read, and a round of serialization and deserialization on each side. That’s why materializing every transformation into a topic is so expensive. I cannot stress this enough, I’ve seen too many times when a Flink pipeline emitted a Kafka topic as a result, only to write it to a datalake (or some other database) with another job. Do you really need that topic?
Intermediate topics are the obvious copies. Eliminate them when possible. The less obvious ones:
- Kafka has no server-side filtering or projection, so every consumer reads everything. Ten consumers that each need 5% of a topic move the whole topic ten times. Read once, route to many: one job that filters and fans out to multiple destinations, or domain-specific substreams instead of a single firehose.
- A mirrored cluster is a full copy of your data. Most mirroring tools also decompress and recompress every batch, so if yours can pass compressed batches through as-is, use it. A REST proxy in front of Kafka is yet another hop.
- Backfilling through Kafka means keeping (or re-producing) history in the most expensive storage you have. Tiered Storage doesn’t completely solve this. Keep history in the lakehouse and use something like Flink’s HybridSource: read the table first, then switch to the topic.
- Hot standbys, blue-green deployments and “one copy of the app per region” all increase compute and reads. Sometimes that’s the right trade-off, but make it deliberately.
But don’t overcorrect! Merging everything into one giant job means a bigger blast radius, shared restarts and shared backpressure. My rule of thumb: split jobs where interference hurts (backpressure, restarts, shared state), and merge them where fixed per-job overhead dominates (JVMs, coordinators, duplicated source reads, monitoring). Finally, keep in mind that adding new topics is always easier than removing existing ones.
Serialization Can Be a Huge Part of Your Bill
Most streaming pipelines are stateless: read, filter, route, enrich, write. For them, the most expensive thing isn’t the business logic. It’s serialization and deserialization. At Shopify, we once found Kryo eating 50%+ of CPU. When I first started using Flink CDC to extract data from Postgres, 60%+ of CPU went to JSON serdes (and the data was serialized twice). Business logic rarely comes close.
Here’s how I approach it, from the most to the least effective:
- If a component just moves data, treat the payload as raw bytes. Put routing and filtering metadata (and schema IDs) in message headers or a message envelope: deserializing the whole message just to skip it is very wasteful.
- Project early and defer deserialization to the operator that actually needs the field.
- Avoid JSON if you can. Avro and Protobuf are faster, smaller and come with strong schema support.
- If you’re stuck with JSON, stop trusting the defaults. Specialized, schema-compiled Avro deserializers were ~3x faster than the standard ones. Measure everything and don’t be scared to try new tools. Also, don’t forget about JSON schemas.
- Never use Kryo in Flink. It silently falls back to Kryo for types it doesn’t support (and there are many unexpected edge cases). Set
pipeline.generic-types: falseso it fails loudly instead, and enable object reuse where it’s safe.
And compress. Compress everything you can! Streaming systems are rarely CPU-bottlenecked, so compression almost always makes a lot of sense. The Java producer’s default compression.type is still none, and JSON compresses extremely well. zstd is a great default: close to gzip’s ratio at a much lower CPU cost.
Of course, I can’t avoid mentioning columnar data! I strongly believe that stream-processing should favour columnar data (think Apache Arrow) from now on. To summarize, columnar data is great for vectorization and offers better CPU cache locality, compression/encoding, and data skipping by design.
Finally, send less if you can. Drop unused fields at the source, use short numeric IDs, and sample noisy events. Sampling is hugely underrated, but it’s a powerful tool, especially if you can control it with feature flags at runtime.
Pay for Latency Only Where You Need It
Cross-AZ traffic is the most famous line item in any classic cloud Kafka bill: replication, plus consumers reading from leaders in other zones. If you haven’t enabled fetch-from-follower yet (broker.rack and replica.selector.class on the brokers, client.rack on the consumers), do it asap.
The real fix is architectural, and WarpStream was the first example of a new architecture: stateless agents, object storage for durability and replication, and clients routed to agents in their own zone. When I ran WarpStream in production, the recipe for paying $0 for inter-AZ networking end-to-end was simple: an agent group in each zone, zone-aware client configuration, and stream processors pinned to a single zone.
The price is latency: around half a second at p99 by default, ~150 ms with S3 Express One Zone. But, in my experience, most pipelines don’t care. Populating a lakehouse, CDC into search or cache, observability, ML features: all fine. Fraud detection, trading, multi-hop user-facing flows: probably not. So stop running everything on the lowest-latency tier. Keep a small low-latency cluster for the workloads that truly need it, and move the rest to a diskless one. You should pay for latency in proportion to how much you actually need it.
State Is a Design-Time Decision
Going from a stateless pipeline to one with hundreds of millions of keys in RocksDB can cost you 30-50x in throughput per core. You can tune RocksDB all you want, but most of that cost is decided upfront, before anyone writes code.
- Beware of join cascades. A chain of binary joins stores intermediate results again and again, so state can grow far beyond the size of the inputs.
- Use pre-aggregated tiles instead of long sliding windows (which keep a copy of each record per window), and sketches like HyperLogLog instead of exact distinct counts.
- Use local disks! It’s the number one advice I give teams that want to optimize large, stateful Flink pipelines. RocksDB on network-attached storage is often IOPS-bound, not CPU-bound. So, forget about EBS volumes and stick to instance-level storage. I’ve seen 5x - 10x difference in performance simply by switching disks.
What About Upserts?
Upserts are a special case: the answer depends on where they happen.
At the edge of a pipeline, they’re great. An upsert-capable sink handles any duplicates, not just the ones caused by producer retries, and it’s often simpler and cheaper than exactly-once delivery. Partial updates can even replace joins entirely: each input writes its own columns into the same table by primary key, with zero join state.
Inside a stream processor, updates can be expensive. In the case of Flink, retractions roughly double the number of records flowing through joins and aggregations, and CDC-style before/after images roughly double the bytes.
So, my rule: if possible, append-only in the middle, upserts at the edge.
Idle Capacity
Fixed clusters are never right-sized: they’re too expensive when traffic is low and too slow when it spikes.
Per-job overhead matters too. Flink excels at stateful computations at scale, but its fixed cost per job is real: two JVMs and a few gigabytes of memory before a single record is processed. That’s worthwhile when a job has hundreds of workers and wasteful when it has two.
Finally, autoscaling. It truly works, and the savings can be dramatic for pipelines with daily traffic cycles. But do it last, not first.
Make Cost Visible
None of this works if you don’t know where the money goes. I find it useful to break the cost of each pipeline into four dimensions: compute (hours), state and checkpoints (bytes and object storage requests), network (bytes), and sink writes.
Find the metric that proves which dimension dominates, and start there. Most teams watch CPU and bytes. Far fewer watch request counts or cross-AZ bytes, and these are exactly the things that are hardest to attribute on a cloud bill.
Enforce client IDs, attribute cost per pipeline, and ship a shared client library (or at least a config) with sane defaults from day one.
Summary
Streaming pipelines get expensive because of their structure, not their volume. So:
- Count requests, not bytes. Batch everything.
- Minimize copies: every hop is paid for several times.
- Treat serialization as your real compute bill.
- Pay for latency only where you need it, and go diskless for the rest.
- Design state on the whiteboard.
- Append-only in the middle, upserts at the edge.
- Get rid of idle capacity. Autoscale last.
- Make cost visible and owned.
RIP, vector database
turbopuffer is abandoning the ANN vector index as its primary storage layout in v3 to eliminate write amplification and improve support for general-purpose queries.
Summary
Deep Dive
- The current ANN-primary architecture forces storage of entire documents under an 'ANN address', causing massive duplication for multi-vector document representations.
- Writes cause cascading rebalances of entire documents when vectors move.
- Vectorized query engines (like those in DuckDB/ClickHouse) require large, uniform blocks, which the current cluster-based ANN index prohibits.
- V3 separates document storage from ANN indexing, allowing each access path to be optimized independently.
Decoder
- ANN (Approximate Nearest Neighbor): A search algorithm that finds data points close to a target vector without scanning every entry, prioritizing speed over absolute accuracy.
- Write Amplification: The phenomenon where one small write operation results in multiple physical write operations to storage, degrading performance and life span.
- Vectorized Execution: A query processing style that performs operations on batches of data rather than individual records, maximizing CPU efficiency.
Original Article
RIP, vector database
We are changing turbopuffer's storage architecture to take search to the next level. turbopuffer v3 changes how documents and indexes are laid out, written, compacted, and queried in turbopuffer. It will allow us to make search faster in every respect — including text, regex, and vector search — but it also lays the foundation to move many more SQL queries to turbopuffer and make them fast.
turbopuffer launched as a serverless vector database (v1), highly specialized to the task of serving extremely cheap and reasonably fast vector searches. Object storage as the source of truth gave the economics, and tiered NVMe SSD/memory caches gave the performance. The value of these particular tradeoffs was validated by our earliest customers, including Cursor and Notion.
turbopuffer evolved to have very strong text and regex search (v2), and is being used for many non-search use cases, like Linear's syncing engine. The query engine has evolved along the way to support all of these query plans, but the storage architecture has remained largely unchanged: the ANN vector index was and still is the primary index around which all other indexes and query plans revolve. This design has constrained several query plans, like GROUP BY and aggregations.
We've pushed the vector-primary architecture as far as we can, and it's time to move on. We're in the process of moving to a new primary index, and making ANN "just another" secondary index. We thought it might be fun to open up the doors and let you follow along.
For this first update, we'll set the stage with why we're doing this in the first place. Walk with me on a short journey from tpuf v1 to today.
v1: an ID and a vector
In the first version of turbopuffer, documents consisted of nothing but an ID and a vector. The prevailing wisdom at the time was graph-based vector indexes, but a hierarchical clustering index plays better with object storage. We started with SPANN, and eventually migrated to SPFresh to support incremental indexing. Vectors are clustered into groups, whose centroids are clustered in turn, repeated to form a tree with a single root.
┌───────────────────┐
│ root centroid │
└───────────────────┘
╱ │ ╲
╱ │ ╲
┌───────────────────┐ ┌───────────────────┐ ┌───────────────────┐
│ leaf centroid │ │ leaf centroid │ │ leaf centroid │
└───────────────────┘ └───────────────────┘ └───────────────────┘
╱ ╲ ╱ ╲ ╱ ╲
╱ ╲ ╱ ╲ ╱ ╲
┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐
│ vector │ │ vector │ │ vector │ │ vector │ │ vector │ │ vector │
└────────┘ └────────┘ └────────┘ └────────┘ └────────┘ └────────┘
We implemented this on top of a storage layer presenting as a key-value map, with sorted and unique keys. Each cluster is given a ClusterId, and vectors within each cluster are given a dense LocalId.
// leaf vectors
K::Vector(C0L0) = vec![0.45, 0.32, ...]
K::Id(C0L0) = 7
K::Vector(C0L1) = vec![-0.28, 0.96, ...]
K::Id(C0L1) = 13
// cluster centroid for C0 is itself clustered at the next level of the tree
K::Vector(C1L4) = vec![0.64, -0.48, ...]
K::Id(C1L4) = C0
As you can see above, everything is keyed by ClusterId and LocalId (e.g. C0L1), which together we call the ANN address. This is what we mean when we say the ANN index is the primary index.
v2: attribute filtering and full-text search
Two new query plans marked the informal transition from turbopuffer v1 → v2: attribute filtering and full-text search.
Attribute filtering
Naturally, customers wanted to be able to add attribute values and filter vector searches on them. To make filtering fast and high-recall, we modeled these as an inverted index that maps an attribute value to the ANN address of the documents that contain it.
K::AttrIndex("family", "Alcidae") -> vec![C0L3, C1L2, C1L3, ...]
K::AttrIndex("genus", "Fratercula") -> vec![C0L3, C1L2, C1L9, ...]
For projections (include_attributes), we also stored the document attributes alongside the ID and the vector.
K::Vector(C0L0) = vec![0.45, 0.32, ...]
K::Id(C0L0) = 7
K::Attr(C0L0, "family") = "Alcidae"
K::Attr(C0L0, "genus") = "Fratercula"
Full-text search
BM25 full-text search was another obvious and much-demanded query plan. Similar to attribute search, full-text search works by first finding the documents that have the query term present (commonly called "postings"). For an FTS index, we also include the (term count, document length) metadata necessary for BM25 scoring:
K::FTS("description", "Atlantic") -> vec![(C0L0, 2, 37), (C9L4, 1, 42), ...]
K::Attr(C0L0, "description") -> "A sharply dressed black-and-white seabird with a \
huge, multicolored bill, the Atlantic Puffin is often \
called the clown of the sea. It breeds in burrows on \
islands in the North Atlantic, and winters at sea."
The problem with a vector primary index
The ANN primary index has largely remained intact until today for one simple reason: it works really, really well for ANN search on object storage. On top of this architecture, we've pushed vector search to single indexes of 100B+ vectors serving 200 ms p99 reads at 1k+ QPS. Any significant change here risks introducing regressions in ANN performance.
However, this layout holds us back from being state-of-the-art for the non-vector query shapes we support, in three main ways: storage amplification, write amplification, and limited vectorization.
Storage amplification
As described above, turbopuffer currently puts the full contents of each document under its ANN address. When there is only one vector, the non-vector data is stored alongside the vector only once.
However, for multi-vector representations of a document, such as document nesting or late interaction, this means we have to duplicate the contents for each vector.
Write amplification
Any time a document is inserted, updated, or deleted, SPFresh may rebalance the vectors to ensure they remain well clustered. Because everything in a document is stored keyed by the ANN address of the document's vector, this rebalancing cascades to moving the full document contents, as well as any inverted (attribute and FTS) indexes that reference it. Updating just one vector can move hundreds of attributes and their indexes.
Limited vectorization
Modern query engines are vectorized: they run tight loops over blocks of values, which amortizes fixed per-block costs, compresses better, keeps the CPU pipeline full, and unlocks SIMD. Every query plan has an optimal block size, but today they are all constrained by the ANN primary index. A plan that wants blocks of thousands of documents to keep the CPU saturated is still stuck at 100–200.
RIP, primary vector index
The solution to these problems is simple: don't key on the ANN address. That is precisely the change turbopuffer v3 makes. As you can imagine, it is not a trivial change.
v3 is a new foundation that will unlock significant performance improvement on all query plans, and we hit a major milestone earlier this month: 100% of CI passes on turbopuffer v3. We started by focusing on correctness. Now we will make it correct and fast. Watching benchmark numbers go down is great fun, so we wanted to get you in at day zero of perf grinding. We will share the benchmarks in public over the coming weeks, as we work toward (and beyond) performance parity before rolling out v3 to production.
Handling hot shards
Slack learned that sharding by a high-level tenant ID creates 'whale' hotspots, necessitating a move toward table-specific keys that match actual access patterns.
Summary
Deep Dive
- Sharding by tenant ID is effective until a specific tenant outgrows the capacity of a single physical server.
- Vertical scaling is a temporary patch; permanent fixes require resharding.
- A 'whale' tenant (large entity) can cause uneven load, breaking the 'even distribution' assumption of naive sharding.
- Sharding by the column queries actually touch (e.g., channel_id instead of workspace_id) keeps hot paths localized to single shards.
- Vitess/Neki allows re-sharding without downtime by streaming writes to the new shard set during migration.
Decoder
- Whale Tenant: A single, exceptionally large customer that consumes a disproportionate amount of infrastructure resources.
- Shard Key: The column used by the database to determine which physical shard a row should reside on.
- VTGate: The routing layer in the Vitess database clustering system that directs queries to the appropriate database shard.
Original Article
Handling hot shards
Since the launch of Neki we've talked a lot about sharding with basic examples to demonstrate how Neki splits a table's rows evenly across shards. Neki's router reads your data topology to handle the placement of writes and to find the correct shard for reads.
But life in production is never so simple.
Imagine your application's Neki database is sharding on a tenant_id column. Makes sense. You want a nice even distribution of tenants across shards. But your app takes off, congratulations! Now, some tenants are hitting their shard at a greater size or volume than most. New product opportunities emerge that require cross-tenant queries.
Suddenly an even distribution of tenants is less useful than an even distribution of data.
You probably just have the wrong shard key, at least for some of your data. No problem, Neki gives you the tools to mitigate increased pressure on shards, and adapt your resharding strategy, without downtime.
A story of resharding success
While Neki is new, sharding isn't. Neki is built by the maintainers of Vitess: a proven, scalable and flexible solution that has a great history of solving these exact problems.
Slack has published multiple articles (1, 2) and talks (3) on how they had to modify their approach to sharding as product demand and requirements changed over time thanks to Slack's rapid rise in popularity.
Their original tidy and obvious way to shard message and channel data was by the ID of the workspace to which they belonged. The logic of which workspace belonged to which shard was maintained in a cluster dedicated to sharding metadata.
The assumption was reasonable, that no single customer would ever outgrow the biggest database. But then a workspace of 10s of 1000s of users lands. Then 100s of 1000s. Vertical scaling and isolation bought time, but not enough.
Over time this became problematic.
With per-workspace sharding, a single hot tenant's messages table quickly overwhelms the shard
As the product got more successful, some workspaces were far more demanding than others. Workspaces for large enterprises could contain over 100,000 users, which ballooned the initial payload the client application needed. Some workspaces' messages tables were getting too large for any one shard. Cross-workspace messaging became a requirement. A new enterprise grid feature needed to organize multiple workspaces under a single enterprise organization. All of this was being made complicated by the current "shard by workspace" strategy.
Their solution involved Vitess, a sharding solution for MySQL, to simplify resharding, starting with sharding messages by channel instead of workspace.
Not every table changed. Each table was explicitly sharded by the column that made the most sense for it: user ID, channel ID, or workspace ID. This created a simpler path to spreading out data load across shards and enabling cross-workspace communication.
Sharding the messages table by channel across shards smoothed out load and unlocked new product opportunities
Explicit vs automatic sharding
Among other benefits, shard allocation was no longer hidden in application logic and a separate metadata cluster. It was defined explicitly in Vitess and enforced by VTGate. Neki's equivalent is the data topology: a declarative configuration that agents and humans can read to reason about where data is written and where it can be read from. You define the columns on which specific tables are sharded and the range of values each shard receives.
This is in contrast to automatic sharding solutions, where the database decides placement for you, which can result in unexpected or unpredictable placement of rows.
Agents love declarative configuration like a data topology because it's foolproof to reason about exactly where data will land and why. Which tables are sharded, which key they're sharded by, and which shards sharded rows are sharded to are all determined by the data topology.
At any time, you can add shards and change sharding strategies. After which, Reshard copies existing rows to new shards as required and streams ongoing writes while the source keeps serving. When you switch traffic, Neki moves reads and writes to the new placement without taking the application offline.
Solving resharding with Neki
Should you suffer from the same success, and have already migrated to Neki, you're in a great position as it gives you the tools to mitigate increased demand and/or adjust your sharding strategy with minimal effort and disruption.
Here's three ways to handle increased demand and requirements. The first two buy you time, the last one is the best long-term solution.
Fix 1: Scale up
If increased demand has your databases hitting their limits, the easy answer is just "add more resources." You could do that and stop here.
Each shard in a Neki database is a distinct Postgres cluster with its own resources. Configuration profiles can be distinct or shared across shards. You can temporarily solve your large tenant problem by assigning its shard a unique profile, vertically scaling it, and going about your business.
On Vitess, per-shard sizing does the same job by giving a hot shard a larger cluster size than the rest of its keyspace.
Fix 2: Isolate large tenants
Additionally, if your larger tenants are causing issues for their neighbours, you can isolate the range of tenants on any one shard.
Here's a visual representation of distributing rows across shards. Note that xxhash doesn't create a perfectly even balance of this small dataset, but would be relatively even over 1000s or more rows.
Since you're in control, distribution can be as broad or fine grained as you like. In a simplified sharding example, an even distribution of hashed IDs across two shards would look like this:
{
"key_ranges": [
{ "shard_uid": "shard-a", "end": "80" },
{ "shard_uid": "shard-b", "start": "80" }
]
}
The start and end ranges in the code example above are only matching the first two characters of a hashed ID. You can go much finer and create a tighter range around the whale's tenant_id.
Isolating the whale's data will require resharding all tenants' data. Reshard copies rows onto new shards, so we cannot re-use the original source shards shard-a and shard-b. With three new shards created, the new layout below describes an updated placement to isolate the hot tenant's data.
{
"key_ranges": [
{ "shard_uid": "shard-c", "end": "80a3f1" },
{ "shard_uid": "whale", "start": "80a3f1", "end": "80a3f2" },
{ "shard_uid": "shard-d", "start": "80a3f2" }
]
}
However, you've now created an environment where one very large tenant can have one extremely large table. A table so large it too would benefit from being sharded. If CPU demands don't get you, storage will. It may be time to make some structural sharding changes.
Fix 3: Reshard some tables
Vertical scaling and isolation of a tenant will only get you so far, but neither of the previous two fixes solves the root cause of your problem. Yesterday's sharding strategy is unsuitable for today's requirements.
Changing sharding strategy doesn't mean sharding or resharding everything. Unless you're Meta or Google you probably don't need to shard your users table.
Investigate your application's access patterns and query shapes to work out what other dimensions data can be sharded by. Look for tables which are regularly joined by a common column key, reshard so they are kept together.
Back to Slack's example, messages were always queried by channel, never by workspace, and channel ID was already part of the messages table's primary key. The right shard key had been there all along. Sharding messages by channel ID made fetching a channel and its messages a single-shard query, even when those messages were authored by users in different workspaces.
Cross-channel queries for messages, which would now be cross-shard queries, were limited to administrative or batch operations, not the critical path. This change cooled off hot spots and gave the team a lot more runway in terms of CPU and storage. Meanwhile, keeping the users table together was critical, as searching for "all users in this workspace" was a common access path.
While resharding isn't something you'll want to do often, it's simpler once you are already within a sharded database. During resharding, Neki will continue to serve queries to existing data until the operation is complete.
Conclusion
Landing big customers and growing your product are good problems to have, but how much stress it causes you depends on the foundation you already have in place. With Neki as that foundation, you are choosing something you and your agents can easily understand, change, and adapt to, no matter which dimension your data grows by.
The longer you wait, the more tables outgrow their original shard key. What could have been one change becomes several at once. Don't wait for a perfect future state. Shard the table that hurts today and address the rest as they need it.
Shard what hurts now.
Parquet X-ray (GitHub Repo)
Parquet X-ray allows developers to inspect Parquet metadata and file structure in the browser without downloading massive datasets.
Summary
Decoder
- Row group: A horizontal partition of data in a Parquet file, typically containing a set of rows.
- Column chunk: The data for a specific column within a row group, stored in a contiguous block.
Original Article
See how a Parquet file is laid out on disk: row groups, column chunks, pages, indexes, bloom filters and the footer.
Paste a Hub URL, an hf:// path, any URL that allows CORS range requests, or open a local file. Only the footer and page indexes are downloaded, so a 2 GB file opens after reading about 3 MB. Parsing uses hyparquet.
Development
npm install
npm run dev # http://localhost:5173/?url=sensors.parquet
npm run verify # lint, type-check, tests, build
See CONTRIBUTING.md for how the code is organized.
License
Apache 2.0
Splink 5: Probabilistic record linkage at billion-row scale
Splink 5 achieves billion-row record linkage in minutes by leveraging the DuckDB 2.0 engine and a new chunked processing architecture.
Summary
Deep Dive
- Scalability: Predict jobs can now be chunked to distribute workloads across multiple machines.
- Dependency footprint: Removed heavy dependencies like Pandas and NumPy, now requiring only sqlglot, duckdb, and pyarrow.
- Incremental linkage: New
predict_within()andpredict_between()APIs avoid full re-runs when adding new data. - Profiling: New profiling APIs for SQL pipelines to identify bottlenecks.
Decoder
- Record linkage: The process of identifying records in different datasets that refer to the same entity.
- Deduplication: Identifying and merging duplicate records within a single dataset.
- Blocking: An optimization technique that restricts record comparisons to only those that share a common attribute (e.g., zip code), significantly reducing the number of total comparisons.
Original Article
Splink 5.0.0 released
Splink is a free and open source library for record linkage and deduplication, capable of processing 1 billion records in less than 10 minutes. It is widely used in government, academia and the private sector and has been downloaded over 22 million times.
We're pleased to release Splink version 5, which is more scalable, faster to train models, lighter to install, and easier to run in production than Splink 4.
Backwards compatibility
There has been no change to the statistical model. Models trained in Splink 4 produce the same results in Splink 5, and the model serialisation format is unchanged, so models saved from Splink 4 in .json format can be loaded directly into Splink 5.
However, Splink 5 syntax is not fully backwards compatible and Splink 4 scripts will need small adjustments to work in Splink 5. Most changes are mechanical, and the core workflow - train a model, predict, cluster - is unchanged.
Main enhancements
- Built for very large jobs, with progress updates.
predict()now supports splitting the work into chunks. This allows you to quickly compute a slice of the overall results table, and also allowspredict()to log progress updates and estimated time to completion. This also enables large jobs to be split across multiple machines when using DuckDB. - Complete large jobs faster. In our tests using the forthcoming DuckDB 2.0 engine, a 1bn row, 10bn comparison
predict()job took 8.5 minutes to complete on a 192vCPU EC2 instance. Significantly larger jobs are possible with chunking. - Faster and simpler training.
- When using a large sample size,
estimate_u_using_random_sampling()is much faster thanks to chunked processing with early stopping once every comparison level has enough observations. - EM training using
estimate_parameters_using_expectation_maximisation()gains amax_pairscap so a loose training rule can be kept while capping the work it generates. estimate_probability_two_random_records_match()has arecord_sample_proportionargument to estimate from a sample of records rather than the full dataset.
- When using a large sample size,
- Faster, easier blocking analysis. Comparison counts are now estimated from a record sample by default, making blocking-rule design much faster on large data. Exact counts remain available with
record_sample_proportion=1.0. In addition to standalone functions, blocking analysis is now available on thelinkerobject for convenience. - Fewer dependencies for simpler and safer installs. Splink now depends on only
sqlglot,duckdbandpyarrow, which themselves have no dependencies. Pandas, NumPy, Altair and Jinja2 are now optional. For pandas inputs or outputs, install pandas separately. This makes Splink quicker and easier to install, reduces dependency conflicts, and substantially shrinks its software supply chain surface. - Incremental linkage is more cleanly supported. If you have already linked a large dataset and receive some new records, it's common to want to create only the new pairwise comparisons, avoiding the need to re-link the entire dataset. This can now be achieved using the new
predict_within()andpredict_between()API. This is a more flexible and robust replacement for the previousfind_matches_to_new_records()function.
Smaller enhancements
Some highlights of other improvements:
- A clearer input contract. Inputs are now registered as
SplinkDataFrames before being passed to theLinker, usingdb_api.register(df, dataset_display_name="..."). Thedb_api=argument has been removed from theLinker- it is derived from the registered data. Source-dataset names forlink_only/link_and_dedupeare set explicitly at registration rather than inferred from positional ordering. - Direct Parquet materialisation. DuckDB can now materialise intermediate and final results directly to Parquet, helping large jobs reduce memory pressure and making it easier to work with outputs that are too large to keep in memory.
- More Pythonic logging. Splink no longer configures Python's root logger, making it easier to embed in larger applications. New
VERBOSE,DEBUG,PIPELINEandSQLlogging levels give finer control over how much detail is shown. - Richer outputs.
SplinkDataFramenow exposesas_record_list(),as_dict(),as_pyarrow_table().SplinkDataFrames now have aquery_sql()method. - A reworked pairwise scoring API.
compare_two_records()is replaced byscore_pair()(one explicit pair) andscore_pairs()(Cartesian product, no blocking). - Match weights instead of Bayes factors. Output tables now contain match weights rather than Bayes factors, with column prefixes changing from
bf_tomw_. This makes results easier to interpret and the algorithms more numerically stable. - SQL pipeline profiling.
DuckDBAPIWithProfilingandSparkAPIWithProfilingare drop-in replacements that write detailed per-query profiles to disk, making it much easier to find the expensive stage of a job. - Cleaner SQL. The SQL generated by Splink is now easier to read. View it by setting logging to 'SQL' level in
Linker(df_sdf, settings, log_level="SQL")
Updating Splink 4 code
Conceptually, there are no major changes in Splink 5. Splink 5 code follows the same steps as Splink 4, uses the same core estimation and prediction routines, and produces the same results for the same settings.
Minor changes to syntax are required to upgrade Splink 4 code to Splink 5. You can find an LLM prompt that should help you automatically upgrade any Splink 4 scripts to Splink 5 here.
Anthropic says it fixed Claude's writing. I ran the evals to check
Testing Claude Opus 5.5 confirms Anthropic significantly reduced 'Claudisms' and em dashes, though the model still relies heavily on specific rhetorical patterns.
Summary
Deep Dive
- Methodology: Used a dataset of 20 frozen research briefs, run across multiple models, to calculate style density per 1,000 words.
- Claudisms defined: Identified as rhetorical 'moves'—signposts, verdict intensifiers, and salience flags—rather than just a vocabulary list.
- Validation: Using an LLM judge (gpt-6-luna) with a 'deletion test'—if deleting a phrase loses no information, it is a Claudism.
- Results: Opus 5.5 is objectively less mannered, though it has traded some tics for others, such as increased bulleted list usage.
Decoder
- Claudisms: The idiosyncratic, often filler-heavy writing style characteristic of Claude models, frequently criticized for being overly dramatic or repetitive.
- Span judge: An LLM-based evaluator that identifies and categorizes specific segments of text rather than assigning a single score to an entire document.
Original Article
Anthropic says Opus 5.5 fixed Claude’s writing. I tested that claim with the standard eval loop in Arize AX: build a dataset, run an experiment for each model, annotate the output by hand, then build an evaluator from the annotations and run it.
The em dash really is gone: 12.9 per 1,000 words in Opus 5, and two in the entire 57,000 words of Opus 5.5 output. The rest of the Claud-isms halved. Better, but not fixed.
I’m sick of Claudisms. You know the ones. You’re scrolling LinkedIn or X, and a post tells you that something “isn’t just a tool, it’s a mindset,” or that a detail “is doing a lot of work,” and that you should sit with that. Humans don’t really write like that, but Claude does, and so does everyone who pastes Claude’s output straight into a post.
When Opus 5.5 came out, one of the launch details caught my eye. The Opus 5.5 announcement says, “We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5,” and that it “is less likely to use jargon or idiosyncratic phrases.” Anthropic staff went a step further. Sholto Douglas posted, “We fixed the writing,” and Tom Brown posted, “we fixed the accent.” The Decoder’s headline said Anthropic “promises less ‘Claudish’ writing.”
That’s a testable claim. And I work at Arize, so when someone says a model got better at something, my first reaction is “show me the eval.” Vibes don’t count, whether they’re mine or a vendor’s.
So I built claude-compare, and ran the same loop I’d use to test any change to an agent: build a dataset, run experiments on it, annotate the output, then build an evaluator and run it.
Step 1: Build a dataset
A dataset is the fixed set of inputs you test against. Every experiment runs over the same examples, so when a score moves, you know it was the thing you changed and not the inputs.
For a writing test, the inputs are writing tasks. I used 20 research briefs, each one containing the notes for a blog post, so that I had long form content to review for Claudisms. The topics cover five genres (opinion, technical explainer, tutorial intro, news analysis, product announcement) and five domains (AI, travel, cooking, games, books). A dataset of AI topics alone would only tell you how Claude writes about AI.
Two things make a dataset like this trustworthy:
It’s frozen. Opus 5 researched each topic once, with web search, and the briefs are committed and hash-locked. The writer gets no tools and no network, so the brief is all it has to go on.
The inputs don’t carry the style you’re testing for. The briefs are written as bullet fragments, so Opus 5’s style can’t leak into the input.
The brief is the load-bearing part of the whole setup (sorry).
Step 2: Run an experiment for each model
An experiment is one run of your task over the whole dataset, with each output stored against the example that produced it. When you compare two experiments, you’re comparing whatever differs between them, so you want that to be one thing.
If you want to compare how two models write, the first step is to just change the model. Sounds obvious, but folks often mix the model change with a prompt change as well. Changing the model first, measuring, and then changing the prompt if necessary, gives you a much clearer view on the impact of your changes.
That means pinning everything else. Both models get an identical prompt. Effort is pinned to medium, because Opus 5 defaults to high and Opus 5.5 to medium, and leaving the defaults in place would have measured two different settings. The harness also checks the served model on every response, so each experiment really is the model it says it is.
Model output varies from run to run, so I ran each model over the dataset twice. That gave four experiments (two models, two repeats), 80 posts and about 110,000 words. In AX the four experiments sit on the same dataset, so every evaluator I add later scores them all the same way and I can compare Opus 5 and Opus 5.5 side by side.
Step 3: Annotate the output
Annotations are human labels on experiment output. They’re your ground truth: what a perfect evaluator would say. You want them before you build the evaluator, because they tell you what you’re actually measuring and they give you something to check the evaluator against later.
You should label at the level you want the evaluator to work at. A score per post would tell me which posts felt Claude-ish, but not why. Marking individual phrases tells me which sentences, and an evaluator that returns phrases can be checked line by line.
The labelling tool doesn’t need to be fancy. I put all 40 Opus 5 posts into one Google Doc and read the lot, all 53,000 words, leaving a “Claudism” comment on every phrase that made me wince. That came to 153 flags across 38 of the 40 posts. The flags then went into AX as annotations on the Opus 5 runs, so they sit next to the output they describe.
Then I looked at what I’d flagged, and it wasn’t what I expected.
I had a hand-written list of the famous Claudisms: “load-bearing”, “delve”, “crucially”, “genuinely,” and so on. It matched just nine of my 153 flags. “Load-bearing,” the phrase everyone jokes about, appears three times in 53,000 words of Opus 5.
What I’d actually flagged were rhetorical moves. The biggest group, about 48 of the 153, was the text telling me something was important instead of showing me why: “The interaction matters,” “a fact worth internalising,” “deserves a moment.” Next were contrast reframes (“It’s a topology, not a genre.”), verdict intensifiers (“the honest answer,” “the whole point”), and signposts that tease an insight instead of giving it (“Here’s the part that surprises people…”).
None of those are fixed wording, so no phrase list will ever catch them reliably. It’s not a phrase problem. It’s a moves problem. 🫠
Step 4: Build an evaluator from the annotations
An evaluator scores every run in an experiment automatically, so you don’t have to read 60,000 words again each time something changes. The annotations shape how you build the evaluator.
Start with the output format. My first evaluator was an LLM judge that gave each post a single 1 to 5 score for how Claude-ish it read. That’s easy to build and almost impossible to check. If it says a post is a 4, which sentences made it a 4? You can’t line a single number up against 153 human flags. So the version I kept is a span judge: it returns every Claudism it finds as an exact quote, in the same shape as my annotations, and I can match each quote against my flags to measure recall directly.
The categories come from the annotations too. The judge sorts each quote into one of six moves taken from my flags: salience flag, contrast reframe, verdict intensifier, signpost, gotcha framing (“the trap is”) and stock metaphor (“load-bearing”, “earns its keep”). And the annotations double as test cases. All three “load-bearing” sentences have to be caught, or the run fails.
A few other choices apply to almost any evaluator:
- Use code when you can. Counting em dashes doesn’t need an LLM, so that’s a plain code evaluator with no wiggle room.
- Don’t let a model grade its own family. The judge is OpenAI’s gpt-6-luna, so Claude isn’t marking Claude’s homework.
- Normalise for length. Opus 5.5 writes about 8% longer, so the judge’s quotes become Claudisms per 1,000 words rather than raw counts.
Both evaluators run in AX against all four experiments.
Step 5: Run it
The em dash really is gone. Opus 5 uses 12.9 em dashes per 1,000 words. Opus 5.5 used two in its entire 57,000 words of output. As Brodie Robertson put it, “The em dashes have been deleted I repeat the em dashes have been deleted.” That result comes from a code evaluator, so there’s no judge involved and no wiggle room.
The other Claudisms halved. The span judge found 4.79 Claudisms per 1,000 words in Opus 5 and 2.38 in Opus 5.5, a 50% drop, and every category went down:
| Category | Example | Opus 5 | Opus 5.5 | Change |
|---|---|---|---|---|
| Salience flag | “This matters.” | 1.67 | 0.93 | −44% |
| Verdict intensifier | “The honest answer is…” | 1.10 | 0.38 | −65% |
| Signpost | “Here’s the part that…” | 0.79 | 0.53 | −33% |
| Contrast reframe | “It’s a topology, not a genre.” | 0.63 | 0.27 | −56% |
| Stock metaphor | “load-bearing”, “earns its keep” | 0.43 | 0.13 | −70% |
| Gotcha framing | “The trap is…” | 0.16 | 0.13 | −19% |
| All Claudisms | 4.79 | 2.38 | −50% |
Claudisms per 1,000 words, from the span judge running in Arize AX.
Salience flags, the “this matters” move, are still the most common Claudism in both models. Gotcha framing is too rare to read much into, with 8 uses against 7. The gap holds in every genre and every domain I tested.
Opus 5.5 also picked up some new habits. A simple scan found it uses about 2.5 times as many bold lead-in bullets as Opus 5, 35% more three-part lists, and writes 8% longer. So some of the old tics were traded in rather than dropped.
My favourite detail: Anthropic’s own prompting guide for Claude Fable 5.1 warns about “mannered prose,” using “this point earns its keep” as its example. Opus 5.5 still wrote that a stand mixer “earns its counter space” and that a searing technique “earns its place.”
So the claim holds up halfway. “Fixed” is too strong. “Much better” is fair.
Iterate on the judge prompt
The 50% figure came from the third version of the judge, not the first. An evaluator is a prompt like any other, and you rarely keep the first draft of a prompt. The annotations from step 3 are what let me iterate on it with numbers instead of gut feel.
The first version said Claudisms fell by only 27%. Moving the judge to gpt-6-luna gave 34%. Recall against my flags was high on both runs, so the judge was finding the Claudisms I’d marked.
Here’s the part that bites people. (Yes, I know, sorry.) Recall only tells you what the judge caught. It says nothing about what else it tagged. So I read a random sample of the spans it found in the Opus 5.5 posts. Plenty of them were ordinary writing. “First, a definition.” got tagged as a signpost. “Strain the context window” got tagged as a stock metaphor. “It applies to all output tokens, not only thinking.” got tagged as a contrast reframe, when it’s just being precise.
Only about 40% of the Opus 5.5 spans in that sample were real Claudisms. (That spot check was done by an AI and is small, so take it as rough). The false positives weren’t random noise. Every writer uses plain transitions and ordinary metaphors, human or model, so they formed a floor under both scores and made the two models look closer than they are.
Sit with that for a minute. (Sorry.)
The third version gave the judge a test to run on every phrase before it tags it: imagine deleting the phrase and reread the sentence. If the post loses information, such as a fact, a number, how something works or what to do, the phrase carries information and it isn’t a Claudism. If nothing is lost, it’s filler, and it counts. Delete “This matters.” and the post says exactly the same thing, so it gets tagged. Delete “not only thinking” from “It applies to all output tokens, not only thinking.” and you lose the point of the sentence, so it doesn’t.
I also added examples of what not to tag in each category and limited stock metaphors to the well-worn ones. Precision on the Opus 5.5 sample went up to about 26 in 30. Recall against my flags on held-out posts dropped from 74% to 64%, and it still caught all three “load-bearing” sentences. I’ll take that trade: an evaluator that misses a few real Claudisms beats one that counts plain English as a Claudism.
Each version ran against the same annotations, so every change came with a number for what it gained and what it cost. I tuned on half the posts and held the other half back to check the result. Without the annotations, the 27% would have looked like a perfectly good answer.
What I’d tell you before you try this
If you’re writing with Claude, hunt your own drafts for the Claudisms, more than the em dash or “delve”: telling the reader something matters, knocking down a claim nobody made, teasing a point instead of making it.
The code, the categories and the judge prompt are all in the claude-compare repo, along with a de-styling tool that uses the same categories as an editing brief.
The bigger lesson applies well beyond Claude. When someone tells you a model got better, don’t trust the vibe, and don’t trust a vendor’s word for it. Build a dataset, run the experiment, annotate the output, build an evaluator from the annotations, and iterate on it against those annotations until you trust its numbers.
For the production-trace version of that loop, see Claude’s hillclimb loop for AI agents: start with production traces.
And this matters. 🙃
Introducing the MySQL Analytical Replica (Parquet + DuckDB)
DBTrail is an open-source tool that enables analytical reporting on MySQL by streaming binlogs into Parquet files, allowing for high-performance querying via DuckDB.
Summary
Deep Dive
- DBTrail captures MySQL binlog data to create consistent analytical snapshots in Parquet format.
- It requires no plugins or agents on the MySQL database, supporting AWS RDS and Aurora.
- The tool uses DuckDB for SQL execution, allowing local analysis of the Parquet files.
- It provides row-level history, enabling 'time-travel' queries to view table state at specific points in time.
- A generated
views.sqlfile provides a consistent interface for querying the latest snapshot. - Benchmarks show it significantly outperforms production MySQL for group-by operations on large datasets.
- The system is designed to avoid the 'pipeline' overhead typically associated with tools like Debezium, Airbyte, or Kafka.
Decoder
- Binlog (Binary Log): A set of files in MySQL that record all changes made to the database, commonly used for replication and point-in-time recovery.
- Parquet: A columnar storage format that is highly optimized for analytical queries, allowing engines to read only the necessary columns rather than entire rows.
- ETL (Extract, Transform, Load): The traditional process of moving data from operational databases into a centralized warehouse for analysis.
- OLTP (Online Transactional Processing): Databases designed for fast, frequent read/write operations of individual records.
Original Article
Your application asks MySQL about now: this order, this customer, this cart. MySQL is very good at that.
Then the business asks about the past:
- Finance: revenue by month, for the last three years?
- Product: who churned after the price change?
- Support and compliance: what did this account look like on 3 March?
Those questions land on the same server as your application, and your primary pays for them. A full scan pushes your working set out of the buffer pool. A long read keeps old row versions alive and purge stalls behind it. A GROUP BY over millions of rows spills temp tables to disk. Row stores answer “now” fast. “Over time” is a different shape of question.
Give those queries a copy they can’t hurt. That is what we are introducing today.
What it is
DBTrail is an open-source analytical replica for MySQL. It keeps a copy of your tables made for reports, as Parquet files (an open format that stores a table column by column, so a report reads only the columns it needs) in a folder or an S3 bucket you own. It follows every change in your MySQL, updates the copy on the schedule you set, and you query it with DuckDB, a free SQL tool that runs on your laptop or a server.
How it works:
- The first copy reads all your tables (with mydumper). This is the one moment DBTrail reads your data directly.
- After that, only the binlog. DBTrail reads it from outside, the way a MySQL replica does. Nothing is installed in MySQL: no plugin, no agent on the database host, no triggers. RDS and Aurora work.
- On your schedule, the copy is updated by folding the changes into a new Parquet snapshot. As often as every 5 minutes, or once a day. Your database is not read again for it, unless a table changes shape (see the limits below).
- You query it with DuckDB. DBTrail gives you a
views.sqlfile with one view per table, so you writeFROM state_shop_ordersand get the table as of the latest snapshot.
No ETL. No scripts or jobs to write to move data out of MySQL. No warehouse to run. One Docker Compose stack, and a folder or a bucket.
A lakehouse without the pipeline
“Lakehouse” sounds like data-engineering jargon, but for a DBA it is easy to picture: your tables as open files in cheap object storage, plus a thin layer that makes those files behave like tables, queried by whatever engine you bring. Every part of your OLTP database is still there. It just stopped living in one process: .ibd files become Parquet in a bucket, mysqld becomes DuckDB, Athena or Spark.
The usual way to build one out of MySQL has three to five moving parts: a CDC tool (Debezium, Airbyte), often a stream (Kafka, Kinesis), a table format (Iceberg, Delta), a catalog (Glue, Nessie), an engine cluster. Some builds skip the stream or the catalog, but it is still a pipeline somebody owns.
DBTrail has one: your MySQL, with nothing installed → DBTrail → plain Parquet in your bucket → any reader. What it gives up is listed at the end of this post, said in the same voice as the strengths.
Querying it
Four steps from nothing to a query:
- Install. You need Docker with Compose. One command starts DBTrail, its web page, and the small MySQL it keeps its own index in.
- Connect your MySQL. Create your login in the page the installer opens and add your server. The page shows the SQL that creates DBTrail’s user.
- Take the first copy and pick a schedule. On Snapshots, press Read database now. Then in Settings, pick how often and press Turn on. Left as it comes, it runs once a day.
- Query it. On Snapshots, press Download, then Download the data. Unpack it and run DuckDB inside the folder:
$ duckdb -init views.sql
-- Loading resources from views.sql
D SELECT c.tier, count(*) AS orders, round(sum(o.total), 2) AS revenue
FROM state_demo_orders o
JOIN state_demo_customers c ON c.id = o.customer_id
GROUP BY c.tier ORDER BY revenue DESC;
┌──────────┬────────┬───────────────┐
│ tier │ orders │ revenue │
│ varchar │ int64 │ decimal(38,2) │
├──────────┼────────┼───────────────┤
│ platinum │ 539848 │ 59426197.40 │
│ gold │ 539752 │ 59353658.54 │
│ silver │ 539056 │ 59295443.48 │
│ bronze │ 538385 │ 59242795.95 │
└──────────┴────────┴───────────────┘
One rule worth knowing early: open the copy through views.sql. Another engine reading the Parquet files directly sees each table as of its last full write and can miss the changes stored beside it.
The numbers, and how we got them
We did not want to publish numbers without the method, so here is both.
The setup. One source: RDS MySQL 8.4 with a TPC-C dataset, 73 million rows, about 10 GB, under a sysbench-tpcc load of 180 transactions per second. Five copies attached to it, one at a time, in 30 to 60 minute windows, measured 18 to 20 September 2026: an RDS read replica, DBTrail 0.84 (refreshing every 5 minutes for this test), ClickHouse fed by Airbyte CDC every 5 minutes, MyDuck Server (binlog into DuckDB), and Redshift zero-ETL. Every default kept. Every manual fix counted as “hands”.
How fresh is it? In a separate measurement under about 340 transactions per second, commit to visible in DuckDB took 2.7 to 13.2 minutes, median 5.9, with a 5-minute schedule. Minutes, not seconds, and not yet repeatable, which is why it is also in the limits below.
Other engines read the same bytes
Because the copy is plain Parquet, it is not locked to DuckDB. We ran the same top-5 query over 55.6 million rows from one snapshot file in DuckDB on a laptop, in ClickHouse with s3() and zero configuration (3.0 s), and in Athena after a Glue crawler (2.2 s). Same answer from all three, no conversion. The caveat from before applies: an outside reader has to apply the change files next to the table, which the DuckDB views do for you.
What else comes from the same capture
The binlog DBTrail reads for the copy also carries every row change, with the row before and after. Starting the day you install it, the same install gives you:
- Row history and undo. See any row before and after a change, and get the SQL that puts it back: the
UPDATEwithout aWHERE, theDELETEthat cascaded. DBTrail generates the SQL; it never runs it for you. - Tables as they were at a moment. Rebuild whole tables as of a past time.
- A row as it was, from your own MySQL client. An optional time-travel port speaks the MySQL protocol, for looking at rows, not for running reports:
mysql> SELECT * FROM speakers WHERE id = 67 AS OF '2026-08-19 09:12:00';
- Ask in plain English. Connect Claude and ask about your changes in words. Every tool is read-only.
Where it stops
A copy you can trust is one whose edges you know.
The scheduled copy is for MySQL and Percona Server 8.0 and 8.4, including RDS and Aurora.
Try it
DBTrail is Apache 2.0. All of it: capture, index, console, recovery, the Parquet you open in DuckDB, and the MCP server that lets an assistant drive it.
- Start here: Quick Start
- What it is, in one page: dbtrail.com/docs
- Source: github.com/dbtrail/dbtrail
How AI Agents Are Reshaping UX
Decagon CEO Jesse Zhang warns that personal AI agents like Meta's Muse are rendering traditional web UX irrelevant, forcing companies to build machine-to-machine commerce interfaces.
Summary
Deep Dive
- Agents perform tasks like purchasing, reordering, and support, necessitating a move toward 'robot-first' UX.
- Traditional UI elements like 'purchase-inducing' flows are 'baggage' for agents that do not render screens.
- Five critical pillars for AI-ready enterprise: Identity verification, procedural rules, negotiation limits, exception handling, and high-speed fraud mitigation.
- Business strategy needs to shift from protecting the 'checkout page' to protecting margins against autonomous, relentless comparison agents.
Decoder
- Machine-to-machine (M2M) interaction: Communication between two automated systems without human intervention, where APIs and protocols replace GUIs.
- Agentic workflow: A computing model where AI agents are given high-level goals and have the autonomy to plan and execute sub-tasks to achieve those goals.
Original Article
As personal AI agents such as Muse introduced by Meta spread, how they might change existing user experience has emerged as a point to watch.
In a recent post on social media X (Twitter), Decagon CEO Jesse Zhang (제시 장) said personal AI assistants will also bring major changes to corporate customer experience. "Amazon blocked Muse. Shopify did the opposite and opened a dedicated payment channel for Muse. Consumer-facing companies need to decide which side to stand on and think about a new UX," he stressed.
If a company is large like Amazon, a strategy of directly controlling the shopping experience and blocking external access could work. Even so, Zhang said, "In the end, Amazon will also back down. It is because what customers want is AI assistant integration."
According to him, use of personal agents such as Muse and Instinct is rising steeply. That is because these agents actually handle tasks. In Meta's case, Meta expects Muse to buy items directly using users' money and provides purchase protection of up to $1,000 per transaction.
Zhang stressed, "Every consumer company will face a customer segment it never considered. It is a robot with a wallet. It is a change as big as the shift from offline to online, and then to mobile apps."
He takes the view that transacting with agents is not something that ends with a single API and one page of documentation.
"If the customer is an agent, the existing UX is just baggage," he said. "Agents skip everything, from product displays and purchase-inducing flows to discount benefits that block cancellations hidden three screens deep. Agents do not look at screens. Rules that were enforced through screens now have to be enforced in other ways."
Zhang thinks agents should also take on that role. He said it is work that requires judgment, not rules. In this regard, he stressed five things.
First is verifying identity and authority. Companies must confirm whether an agent truly represents the customer and define the scope of its authority. Browsing and purchasing are different, and reordering $40 and booking $4,000 are also different.
Second is rules and procedures. Handling differs for flight changes, card reissuance and plan changes depending on customer type, laws and internal policies.
Third is negotiation limits. "If you cannot decide who gets a discount and when to stick to principles, agents will secure the terms most unfavorable to the company every time," Zhang said. "Agents do not get tired or embarrassed, and compare competitor offers in real time."
Fourth is exception handling. Even for requests past the refund deadline, a customer of 6 years and a customer who joined yesterday should be viewed differently.
Fifth is fraud happening at machine speed. While legitimate agents may buy goods in 0.2 seconds, malicious agents probe for policy loopholes 10,000 times an hour, so limiting request frequency alone cannot stop them.
The end state Zhang envisions is a structure in which AI customers and AI agents interact in transactions. People set what they want, such as "buy it at the lowest price," and companies set direction by saying, "protect margins and retain loyal customers." Transactions are handled machine to machine.
"A company AI that deals with customers must be able to deal with both people and AI assistants. If it cannot adapt, it will fall behind," he said. "If the effort required for a single conversation falls close to zero, conversations between consumers and companies will increase sharply, and economic potential that has not been realized will also come alive."
The Box Design System: AI‑Ready Figma Files for Enterprise Software
The 'Box Design System' organizes Figma files into bounded, behavior-focused sections to provide AI coding agents with the specific context needed for implementation.
Summary
Deep Dive
- Design file organization is no longer an internal cleanup task; it is a critical input for AI development.
- Bounded boxes provide a focused unit of work for agents, preventing them from hallucinating context from unrelated canvas areas.
- Hierarchical structure: Feature (capability) > Flow (interaction) > Screen-state (moment).
- Arrows and labels are mandatory for AI clarity, acting as a visual map for state transitions.
- Design boxes do not replace code, but they reduce the amount of inference an agent must perform, thereby reducing implementation errors.
Decoder
- Figma MCP server: A connector that allows AI agents to read, query, and interpret Figma file data directly through the Model Context Protocol.
Original Article
On a recent enterprise software project, we were working with a Figma file that needed to support both traditional developer handoff and AI-assisted development. One section alone contained dozens of screens documenting a single part of the product.
Each screen made sense on its own. The difficulty was understanding how they belonged together.
For the designers and developers who had worked on the project for months, much of that context was already familiar. We knew which screens represented the starting point, what action caused each change, which states were alternatives, and which details were still under discussion.
A coding agent coming to the file had none of that background. An individual screen did not provide enough context to explain the complete behavior, while the full canvas introduced too much unrelated information. We needed a way to organize the relevant frames, states, and decisions around a single interaction so they could be understood and used more effectively in AI-assisted development. That need became the basis of our Box Design System.
What is the Box Design System?
The Box Design System is our method for organizing complex product behavior into bounded sections of the Figma canvas. Each box contains the screens, states, transitions, annotations, and implementation context needed to understand one part of the product. This makes the file easier for people to navigate while also preparing its structure for AI-assisted development.
In the enterprise platform we were designing, navigation alone included several connected behaviors. Users could interact with a floating bar, select assets, explore the map directly, or move through a hierarchy. All of these actions occurred within the same broader interface, but they did not belong to the same flows.
Placing every screen in one large sequence would have made the file look complete while leaving its logic ambiguous. Instead, we grouped related states into separate boxes such as “Floating Bar,” “Select Assets,” “Direct Map Exploration,” and “Hierarchy.” Each box represented one coherent part of the experience.
At first, this was a way to make a complicated canvas easier for people to navigate. As AI agents became part of the development workflow, the boxes gained another purpose: they created clearer units of context for implementation.
Here are four principles that shaped the system.
1. Organize boxes around behavior, not pages
A page is a visual container. It is not always a useful development boundary.
One page may contain several independent interactions. A single interaction may also extend across panels, overlays, menus, and several screen states.
In our project, selecting an asset and exploring the map took place within the same primary interface. Visually, the screens shared most of their structure. Functionally, however, they represented different behaviors.
Treating them as one group would require a developer or agent to determine which changes belonged to asset selection and which belonged to map exploration. That distinction was obvious to the team because we had discussed it repeatedly. It was not necessarily visible in the pixels.
We therefore organized the boxes as a hierarchy rather than treating every box as an equal unit. The system has three levels:
- Feature boxes represent a complete product capability, such as Navigation.
- Flow boxes sit inside the feature box and separate interactions such as “Select Assets” and “Direct Map Exploration.”
- Screen-state boxes sit inside each flow and name the specific action or state shown on every screen.
This structure allows someone to move from the broader feature to an individual flow and then to the exact screen behavior without losing the relationship between them. It also means that an implementation task can be scoped around the behavior the team needs to build, rather than every element visible on the page.
2. Make each box understandable on its own
A useful box should not require a guided tour from the designer who created it.
Someone opening the file should be able to understand what the interaction does, where it begins, and what changes as the user moves through it.
That does not mean duplicating every piece of product documentation inside Figma. It means including enough context to remove the most consequential ambiguity.
Depending on the feature, a box may contain:
- The default or entry state
- The action available to the user
- The resulting state
- Expanded and collapsed variations
- Selected, disabled, empty, loading, or error states
- Short notes explaining behavior that cannot be understood visually
- References to existing components or patterns
This self-contained structure matters for human teams, especially when people join a project late or return to a feature after several weeks. It matters even more for coding agents.
An agent does not have the memory of the workshop where the interaction was discussed. It does not know which Slack message changed the requirement or which visual difference was intentional. It can only work with the context it is given.
If critical information sits somewhere else on the canvas, remains buried in a comment thread, or exists only in the team’s memory, the agent must infer what is missing.
And inference is where avoidable mistakes begin.
3. Make interactions and transitions visible
Individual screens show different interface states, but they do not always explain how the user moves between them.
In our design files, we use simple arrows to connect an interaction with the screen or state it produces. If clicking a button opens a sidebar, we connect that action to the designed sidebar state. If an interaction takes the user to another screen, the arrow shows exactly where they will arrive.
Not every interaction needs this treatment. Adding arrows to simple or obvious actions can create unnecessary visual noise. They become particularly useful when one screen contains several possible interactions, each leading to a different state or part of the flow.
For someone who has worked on the product for months, these relationships may already feel obvious. Developers and AI agents do not necessarily have that context. Without a visible connection, an agent may recognize all the individual screens but still misunderstand which action produces which result.
Combined with the box hierarchy, these connections make both the structure and sequence of the experience easier to follow. The boxes show which screens belong to the same feature and flow, while the arrows explain how users move between them. It is a simple practice, but it reduces how much developers and agents need to infer from the designs.
4. Give the AI agent a bounded unit of context
AI coding tools can now access much more than a screenshot. When connected through tools such as the Figma MCP server, an agent can retrieve structured design information, visual references, variables, components, and layout data from a selected part of a file. However, in AI-assisted development, access to more information does not automatically lead to a better understanding of what needs to be built.
An isolated frame may leave out the surrounding states required to understand a behavior. At the other extreme, a large page or full canvas may introduce unrelated screens, unresolved concepts, and several versions of the same feature. The Box Design System creates a practical middle layer between these two extremes.
Instead of asking an agent to “build this page,” a developer can direct it toward a bounded interaction containing the relevant states, transitions, and supporting context. This gives the agent enough information to understand the intended behavior without requiring it to interpret the entire product.
The boundary of a box does not need to match a single code component. Some boxes may translate into one component, while others may involve several components and application states. The box is there to define the product behavior being implemented, not to dictate the final code architecture.
The quality of the generated code still depends on the codebase, prompts, design-system implementation, and the agent’s access to existing components. Figma also recommends semantic naming, variables, Auto Layout, reusable components, annotations, and Code Connect to improve consistency. The boxes complement this infrastructure by organizing those elements into a coherent piece of product behavior that the agent can understand and act on.
What boxes cannot solve
A clean Figma canvas cannot compensate for an undefined product.
If the team has not decided how an interaction should behave, placing its screens inside a box will not resolve the underlying uncertainty. It may simply make that uncertainty easier to see.
That is still valuable.
When related states are placed together, inconsistencies become more visible. The team may discover two competing versions of the same behavior, an exception with no recovery path, or a state that was discussed but never designed.
Boxes also do not replace the platform’s design system. While the main design file uses boxes to organize screens, flows, and product behavior, the design system lives in a separate Figma file containing the shared components, variants, and variables used across those designs.
The design system does not need to follow a box-design hierarchy, but it should be structured and detailed enough for agents to understand exactly what they need and where to find it in the file. Clear naming, logical grouping, and consistent component structures help both developers and AI agents find and reuse the correct elements. Without that structure and the corresponding code, an agent may understand the intended behavior while still selecting the wrong component or creating unnecessary one-off elements.
The box defines the scope and behavior. The design system defines how the interface should be constructed. The codebase defines how it must function in production.
Why Figma file structure now matters for AI-assisted development
For years, design-file organization was often treated as an internal concern. A tidy file was easier to navigate, easier to hand off, and kinder to the next designer who had to work in it.
That is no longer the full story.
As design files become direct inputs to AI-assisted development, their structure begins to influence what an agent understands, what it overlooks, and how much it must infer.
The canvas is no longer only a place where the team presents the final interface. It is becoming part of the system through which software gets built.
Our box-based approach began with a practical need: make a complicated enterprise product easier to understand. Its value grew when the same structure helped divide that product into clearer units for implementation and review.
The boxes themselves are not the innovation. The clarity they create is.
Components define what the interface is made from. Frames capture individual moments. Boxes explain how those moments belong together and what the team is trying to build.
In a development process increasingly shared with AI agents, that context is no longer optional documentation.
It is part of the product.
amazing-glass (Website)
Amazing Glass implements physically accurate refraction for web components, avoiding the common "blurred rectangle" trap by using Snell's Law.
Summary
Deep Dive
- Implements real-time ray bending using Snell's Law to simulate thickness and refraction.
- Uses custom Chromium-based filters to render edge-based light bending.
- Provides interactive controls that adjust thickness, frost, and rim width.
- Includes typed wrappers for React and Vue components.
- Features a design system where components like menus and buttons adaptively morph based on user interaction.
Decoder
- Snell's Law: A formula used to describe the relationship between the angles of incidence and refraction when light passes through boundaries between different media.
- Backdrop-filter: A CSS property that applies graphical effects like blur or color shifting to the area behind an element.
Original Article
Blur is not glass.
Most "glass" on the web is a blurred rectangle. Real glass has a thickness. Light hits the curved rim, bends, and drags the background along with it. Watch the stripes at the edges.
backdrop-filter: blur(14px) A frosted rectangle. Flat edges, no depth.
<ag-glass> A slab with a curved rim. The edge bends what is behind it.
Light bends at the edge.
For every pixel we take the slope of the rim, bend the view ray with Snell's law, and store how far it moved. Chromium runs that map as a filter on whatever is behind the element. Three passes with slightly different strength, one per colour, give the thin rainbow at the rim.
Cross-section of the rim. Rays enter the curved edge and land further in. Thicker glass, bigger bend.
These are the real components. The sliders drive the same parameters you get in code.
We checked our work against Apple's.
A small SwiftUI app draws Apple's .glassEffect over the same pictures, at the same sizes. We screenshot both, compare every pixel, and let a search nudge our parameters until the difference stops shrinking. Drag the line.
Mean absolute error per colour channel, 0 to 255 scale, over each glass shape plus a 10 px margin, two scenes, macOS 27. The biggest surprise: Apple's Clear glass is frosted, about 16 px of blur. We had it at 1.6.
Controls that notice your finger.
Knobs are solid at rest and turn into lenses while you hold them. Menus grow out of their buttons. Selections stretch on the way to where you tapped. Go on, touch everything.
Looks right in an actual app.
Scroll the feed and the tab bar shrinks to get out of the way. Drag across it and the selection follows as a lens. Tap an album and a sheet floats up, inset so the content peeks around it. Pull it up and it goes flush and more opaque.
Corners are concentric. The sheet sits 8 px inside a 55 px device corner, so its corner is 47 px.
Two drops, one glass.
Bring glass shapes close and they melt into one surface, with one rim and one refraction. Move your pointer over it, or drag a finger sideways across it.
Two lines to try it.
One CSS file, one script. Every element works in any framework. React and Vue get typed wrappers with the usual props, events and v-model.
A Short Update on Designing for Foldable Devices
Despite the release of the iPhone Duo, browser support for foldable device APIs remains locked within Chromium, leaving Safari users behind.
Summary
Decoder
- Device Posture API: A browser API that reports whether a device is in a "folded" or "flat" state.
- Viewport Segments API: A browser API that provides the geometry and coordinates of individual screens on a foldable device, allowing developers to manage content placement around the hinge.
Original Article
A Short Update on Designing for Foldable Devices
And here we are three years after I wrote about the Google Pixel Fold being announced, followed now with the announcement of the iPhone Duo...(I have questions about the naming by the way).
Four years ago I was talking about web primitives in the platform for the Surface Duo. My how time flies.
There are CSS media features, a Viewport Segments API, a Device Posture API but Chromium based browsers are the only ones currently supporting these things.
I haven't been able to find any signal yet on whether Safari will support these things in the web platform as the developer docs focus on application development.
If you're interested in trying out the platform features, you can emulate the Surface Duo and Galaxy Z Fold in the developer tools.
And if you're thinking, do I really have to have my website adapt to two screens? The answer is no. Adding a design to an application or dual screen makes sense if you have an experience that has two simulataneous contexts that are useful e.g. a list of email messages/inbox on one screen, an open message, email thread or email composer on the other.
Here's one of my talks from 2022 if you're interested in learning more about what's available in the browser for dual screen/foldable devices.
Happy building :)
Most powerful obesity drug yet: People lost up to 25% of weight in trial
Eli Lilly's experimental drug retatrutide yielded up to 25% average weight loss in late-stage trials, setting a new benchmark for obesity treatments.
Summary
Decoder
- GLP-1/GIP/Glucagon: Hormones that regulate insulin secretion, stomach emptying, and feelings of fullness; the drug mimics these to alter metabolic response to food.
- Prediabetes: A condition where blood sugar levels are higher than normal but not yet high enough for a type 2 diabetes diagnosis.
Original Article
Researchers published late-stage clinical trial data today for the latest obesity drug, retatrutide—expected to be the most powerful formula yet—and the results appear in line with high expectations. Patients with obesity on the highest retatrutide dose lost an average of 25 percent of their body weight after 80 weeks, and an average of 30 percent after an extension period to 104 weeks. Overall, more than a third of participants taking retatrutide lost 30 percent or more of their weight.
The drug also proved effective at reducing knee pain (by up to 62 percent) in a subset of participants with obesity-linked knee osteoarthritis. It reduced the number of sleep apnea events per hour (by up to 57 percent) in a subset of participants with obesity-linked obstructive sleep apnea. The drug improved cardiometabolic measurements across the board, including blood pressure, triglycerides, and low-density lipoprotein cholesterol (bad cholesterol). At the start, more than a third of trial participants had prediabetes, and, by the end, the condition had resolved in more than 90 percent of those participants.
The trial began in 2023 and included 2,339 participants from 131 clinical trial sites in 11 countries. Participants were broken into four nearly equal groups, given either: 4 mg of retatrutide, 9 mg, 12 mg, or a placebo. Across all groups, the average starting weight was around 113 kg (250 pounds), and the average body mass index (BMI) was 40. For the subset analyses, 574 participants had knee osteoarthritis, and 243 had obstructive sleep apnea. The results appeared today in the New England Journal of Medicine.
The data is likely to only ratchet up the anticipation for the drug among the millions of Americans with obesity. Retatrutide has been so sought after that, in June, news broke that top health officials in the Trump administration were involved in granting special access to the experimental medicine to a mystery 79-year-old in April—widely speculated to be Trump, who was 79 at the time.
Next-generation obesity drug
Retatrutide is being developed by pharmaceutical giant Eli Lilly, which also makes tirzepatide, a potent dual-acting treatment for obesity (Zepbound) and type 2 diabetes (Mounjaro). Both medications build on GLP-1-based obesity drug semaglutide (Ozempic/Wegovy) from pharmaceutical company Novo Nordisk. Tirzepatide targets not just GLP-1 (aka Glucagon-like peptide-1), but also GIP (aka glucose-dependent insulinotropic polypeptide).
Retatrutide builds on these drugs by adding glucagon to the combination of GLP-1 and GIP. All three are hormones that have overlapping roles in responses to food, helping the body regulate blood sugar levels, the feeling of fullness, how much we eat, and downstream metabolic factors.
The interplay among the three hormones is complex, and they have different activities in different places in the body and at different times. GIP is secreted by cells (K cells) in the first sections of the small intestine, right after the stomach. This hormone is best known for stimulating insulin release in response to glucose (sugar). But GIP has other activities, including triggering the degradation of triglycerides (a type of fat in the blood) and working in the brain to trigger the feeling of being full. It also has a role when blood glucose levels get too low. In that case, GIP stimulates an increase in glucagon.
Glucagon is a hormone released from alpha cells in the pancreas, and it acts counter to some of the main roles of GIP and GLP-1. Glucagon is best known for triggering the release of glucose and fatty acids into the blood, which it does when blood sugar levels get too low (hypoglycemia). But after meals, glucagon also seems to play a role in increasing insulin production, delaying stomach emptying, and regulating lipid levels.
GLP-1, the most famous of the hormones, is produced by cells (L cells) further down the gastrointestinal tract, namely at the end of the small intestine (the ileum) and the colon. GLP-1 works in response to sugar to increase the release of insulin. It delays stomach emptying and works in the brain to trigger the feeling of being full. It can also send signals to spur the degradation of lipids in fat tissue. Additionally, both GLP-1 and GIP can stimulate the release of another hormone, called adiponectin, which can help with insulin sensitivity and reduce inflammation.
Trial limitations
Retatrutide can stand in for all three hormones (GLP-1, GIP, and glucagon), but its active ingredient is a single synthetic peptide. Different sections of the molecule interact with the specific receptors for each of the three hormones, playing the role of the hormones. However, it doesn’t bind to these receptors with the same strength as the natural hormones; compared to GLP-1 and glucagon, retatrutide is less potent, but compared to GIP, the drug is nearly nine times more potent. The peptide is also stabilized, so it stays active in the blood for six days, allowing for weekly injections.
The effects of the triple-acting drug seem more potent than in the previous obesity drugs. But a notable limitation of the trial is that it didn’t compare retatrutide to an existing treatment, such as tirzepatide. How it compares in weight loss and other health benefits will need to be explored in future trials.
So far, retatrutide’s safety looks similar to what’s seen with existing GLP-1 and GLP-1/GIP drugs: The most common complaints were transient mild-to-moderate gastrointestinal symptoms, such as nausea, diarrhea, and constipation. Some participants reported feeling dizzy and having low blood pressure on retatrutide, which was more common among people taking medication for high blood pressure. Ten people in the trial had cardiovascular events (including one in the placebo group), and there were also 10 reports of pancreatitis (two in the placebo group), but the study wasn’t big enough to assess the risk of these conditions. Additional larger, longer trials are needed to evaluate these rarer potential safety risks.
Commit Description as a Thinking Tool
Writing your own commit messages is a vital cognitive check to ensure you actually understand the changes an AI agent generated.
Summary
Decoder
- Agentic coding: The use of autonomous AI agents that can plan, write, and execute code changes with minimal direct human guidance.
- Commit message: A text description accompanying a code change in version control systems like Git that explains what changed and why.
Original Article
Before the AI era, I wrote pretty long commit descriptions (or commit bodies) for major changes. It took me about five to ten minutes to draft and reread them to make sure I didn’t miss anything important. I did that for several reasons.
I wanted to include all the useful information so readers wouldn’t have to hunt for it in multiple places.
I wanted to explain what and, more importantly, why. The “what” summarizes the changes that are generally self-explanatory, but it gives a starting point for explaining the “why.”
Sometimes, I write it in first person, like I am drafting a message for someone: “I did this because…”, “I am doing this until we…” etc. I then go ahead and explain why, so that it is easier for others, and most importantly for my future self, to understand why we made that change.
It was a good exercise. It wasn’t just about writing the commit message and description. The writing process itself helps me reflect on the code I wrote. I reread the code and summarize the changes. During that process, I tend to re-evaluate the decisions, and sometimes that leads to a different or better change.
Now we are in the era of agentic coding, where everything from code to commit descriptions is written by AI. There is a huge debate on whether we should read the AI-written code, and how difficult that is in terms of readability. The part I find difficult is reading and understanding AI-written commit descriptions.
Agents can write commit messages for the changes they make. But they may not have the full context that is spread across different communication and project management tools. Some of those might be offline too. When the AI doesn’t know the ‘why’ part, it comes up with its own reasoning. I find that dangerous. When we read that later, it may not make sense, because the real reason was completely different.
One obvious solution is to give the agent all the context it needs, through chat or tools. This helps to fix the issues with fabricated reasoning. The agent can now explain the “why” clearly.
But this doesn’t fix the other problem. The agent with the right context will write a convincing commit message. But only I can verify if the code does what the description says. So, here is something I do:
Write the commit message and description myself.
Here is why.
Writing a commit description myself helps me to reflect on the changes AI made. It is also a way to check if everything is as intended. If I cannot explain “why,” I am shipping something I don’t understand, which will be hard to explain or fix if it breaks later. It goes back to the old quote. If you cannot explain it, you did not understand it. The commit description again works here as a thinking tool.
Some parts, like “I am doing this until we…”, are temporary decisions with exit criteria. We sometimes set exit conditions, and AI cannot infer them from code or other tools because they are usually not written down anywhere since they seem too obvious to mention. But writing the commit message forces me to complete that sentence, and it helps future readers decide whether to keep that change.
The agent can write the code and the description. But writing why is where you find out whether you understand what you are shipping.
Tech CEOs Privately Questioned Amodei for Sounding AI Alarm Bells
Tech CEOs confronted Anthropic CEO Dario Amodei for his public warnings about AI safety, questioning why he emphasizes extreme risks.
Summary
Original Article
AI chief executives gathered with Trump at the White House on Tuesday to discuss a statement of principles to keep AI models safe. Behind the scenes, the executives questioned Anthropic CEO Dario Amodei about his warnings about the dangers of AI, asking why he was being so extreme in public about the capabilities of AI models and the risks they pose. Amodei told them it was important to be honest with the public about model capabilities and not play down risks. He believes that safety is possible if the industry and administration work together.
DoorDash launches an AI agent you can text to order food
DoorDash is launching a text-based AI agent that allows users to place complex orders through Apple Messages, moving to compete with Uber Eats.
Summary
Original Article
DoorDash launches an AI agent you can text to order food
DoorDash announced on Wednesday that it’s launching a text-to-order AI agent that lets users place orders through Apple Messages. The new tool allows users to send prompts like “order my usual,” and the agent will understand that they mean their Friday night order.
DoorDash says users can also ask for a specific dish and request a local recommendation. The agent will then search local spots and suggest a cart based on the prompt, and even text photos of the food it recommends. When ordering for a group, DoorDash says the agent can handle mixed dietary preferences and different quantities in one order.
By launching an AI agent for food ordering, DoorDash is looking to gain an edge over rivals Uber Eats and Grubhub. The launch also comes amid a growing push to adopt personal AI agents that can handle tasks without requiring users to navigate apps or websites themselves.
DoorDash is opening up a waitlist for users in the U.S. to try out the new feature.
In addition to the new AI agent, DoorDash announced that it will begin testing its delivery drones with select restaurants in Northern California.
State of Markets II
Tech contributed 76% of S&P 500 earnings growth in 2026, signaling a fundamental shift where software is no longer a sector but the engine of the entire economy.
Summary
Deep Dive
- Tech accounted for 76% of S&P 500 earnings growth in 2026.
- A rotation from pure software to hardware (semiconductors, power, networking) is underway due to AI infrastructure needs.
- Hyperscaler free cash flow is heavily subsidizing semiconductor production.
- GPU demand remains robust; A100 residual values and rental rates are not collapsing as bears predicted.
- Software companies have shifted from growth-at-all-costs to 75% profitability, though only 30% are growing faster than 20%.
- Enterprise AI adoption is broad but remains shallow, with only 2% of companies tracking specific metrics for AI impact.
Decoder
- Bits to Atoms: The industry shift from prioritizing software (bits) to physical hardware, infrastructure, and industrial robotics (atoms).
- ZIRP (Zero Interest Rate Policy): The era of near-zero interest rates that encouraged high-growth, loss-making software startups.
- Hyperscaler: Large cloud service providers (e.g., AWS, Microsoft Azure, Google Cloud) with massive infrastructure footprints.
Original Article
State of Markets II
We are pleased to release the second edition of a16z’s State of Markets.
If you recall the first State of Markets (and even if you don’t), then you know what to expect for SoM II: the canonical chartapalooza of 2026 from the perspective of equity markets and tech (for the first two quarters, at least).
View the charts
Behold, just a small preview of what’s inside.
Tech Is The Everything Cycle
As a force in capital markets, tech really took off following the GFC. After the housing bust and credit crunch, tech offered a high-growth, capital-light alternative to the asset-wary. Perhaps more importantly, tech offered massive embedded operating leverage, given substantially untapped TAM, and software’s near-zero marginal cost. Hindsight is 20/20, but the investors who remained skeptical that loss-making techcos would ever grow into high-margin machines, missed out on the theme of the previous decade.
That’s all true, and software did, in fact, eat the world, but tech’s centrality to capital markets has risen to a whole other level:
Tech has steadily compounded earnings growth, albeit off a lower base, much more so relative to the field. Since 2023, however, it’s fair to say that tech is the earnings growth story, contributing ~76% of the SP500’s total earnings growth in 2026 (as of late August).
What lies beneath that earnings growth is an interesting story in and of itself, but you can’t really call tech just a sector anymore. In the old days, durable goods defined the cycle–houses, dishwashers, cars, etc.–but tech has taken their place. Tech is everywhere and in everything. Tech is the everything cycle, now.
From Bits to Atoms
Within tech, the theme of the year (and beyond), has been the rotation from bits to atoms.
While software dominated the previous tech cycle, hardware is now the star of the show, and it’s pretty easy to see why:
The AI buildout has caused a surge in demand for traditionally sleepy, cyclical, and capital intensive industries like semiconductors (but power and networking, as well). That demand has been funded in large part by the historically massive profits generated by the world’s largest tech companies–functionally converting hyperscaler free cashflow into semiconductor free cashflow—but increasingly by debt, as well.
It’s not just AI that’s spawned a revival for the world of atoms, however. Global infrastructure needs are measured in trillions, defense spending is rising, grids are grappling with the electrification of everything, manufacturing is being reshored, and robots and robotaxis are inbound.
In all events, hardware and infrastructure have emerged as the apple of the market’s eye, after years of trailing in software’s wake. Both public and private equity is funding innovation in compute, memory, power, robotics, manufacturing, and defense, with an intensity we haven’t seen in decades (if not longer). The point being: atoms are so back.
Reports of GPU Obsolescence Have Been Greatly Exaggerated
On the subject of all that capex, one thing has been clear: demand for compute is still outpacing supply.
At least one prominent bear of the buildout expressed some skepticism whether all this money invested in GPUs could possibly be worth it, given that they’d become obsolete in 3-4 years’ time. It’s great that Nvidia is selling B200s like gangbusters, but what does that mean for all the A100s that were installed a year or two before?
Well, for now at least, it turns out that the upward inflecting demand curve for AI-compute has meant that the A100s are still pretty useful:
Rental rates (and thereby residual values for GPUs) are supposed to decrease over time, but that’s not really what’s been happening. As intelligence gets cheaper, demand for compute is only rising, causing GPU pricing to climb upwards for the latest chips (and stay pretty for the older ones). Even the A100 is pricing at-or-above what it was at the beginning of the year.
The story is far from over, of course, but so far, neither compute nor model advancements have been a zero-sum game. Better, cheaper, intelligence is accruing value to the entire ecosystem, with both older chips and models retaining substantial value well past the expiration dates assigned by the bears.
And, oh, by the way, all of this is happening while AI adoption remains relatively immature. Adoption is broad, but for the most part, it’s relatively shallow:
While nearly 30% of SP500 companies report some “quantifiable impact” of AI, only ~2% are reporting any tracked metric. The same goes for agentic use-cases, where a tiny share of the overall userbase is meaningfully deploying agents with any scale.
Likewise, on the consumer side, paid penetration is still tiny:
As of April, barely ~2% of US households were paying for some AI service. The number is higher now, and growing, but it’s still tiny in the big scheme of things.
The point is that GPUs are already running hot, even while the data indicates that it’s still so early when it comes to mature AI adoption and utilization.
SaaS-Prove-It, Not -Pocalypse
The year started with the declaration that software was dead, a sure fire victim of AI’s vibe-coded everything. SaaSpocalypse was nigh. Run for the hills, enterprise SaaS, and don’t let the door hit you on the way out.
Reality, as it often is, turns out to be a bit more nuanced than that. There was surely a sell-off, albeit a more discerning one than “pocalypse” would imply, and while AI (or the threat of AI) is definitely playing a role, it’s not the only thing.
The actual thing about public software is that some part of this reckoning, or really a re-rating, was a long time coming:
Since the end of ZIRP, techcos (but not just techcos) traded growth for profitability. Back in ‘22, the market was full of high-growth, mostly unprofitable software businesses, but by 2026, the story has inverted: ~75% are profitable, but only ~30% are growing 20%+.
It makes sense, insofar as higher interest rates are intended to make capital relatively scarce, and so companies wisely pivoted from loss-making growth, to a more self-sustainable, slower-and-steadier growth. That’s all well and good, but slower-growth companies do not beget high-growth multiples, at least not for long, and eventually the new “slower-growth” normal caught up to software.
Not for everyone, of course. Fast-growing companies are still trading at multiples in line with historical averages (albeit not ZIRP peaks), but there are simply fewer of those now, so the sector as a whole was broadly repriced.
There’s been no apocalypse for software, but there has definitely been a “prove it.”
Looking Ahead At Where Things May Go
And finally, while the lion’s share of the presentation covers what’s been, we’d be remiss if we didn’t offer some thoughts on where we think things may be going:
We expect that AI will expand the surface area of demand. That means advancing towards more mature adoption across enterprise and consumer, but also pushing out into entirely new frontiers in robotics, biotech, health and AD.
In general, the technology is improving at an exponential pace, and while no one can predict the future, given the rate and pace of change, we’re pretty confident that this cycle isn’t going to be like any previous cycle.
View the full presentation
The AI Safety community is unfortunately doing more harm than good
The constant stream of 'AI doom' rhetoric from frontier lab researchers risks fueling public distrust and inviting heavy-handed government regulation.
Summary
Decoder
- AI Alignment: The process of ensuring AI systems act according to human intent and ethical values.
- P(doom): A term used to describe the subjective probability an individual assigns to the risk of human extinction caused by AI.
Original Article
The AI Safety community is unfortunately doing more harm than good
Just yesterday, Palisade Research published a series of interviews with 22 active or former employees of frontier AI labs, asking them to share their thoughts on the topic of AI doom and relevant scenarios like human extinction, or what it would actually take for a computer to kill all of us.
The current and former employees share their opinions in these high quality interview clips, stating how "human extinction is a coin flip" or that "humans are in the way of these models" in an effort to, I assume, raise concern over continued AI development that they believe threatens humanity.
This would have come as a shock, or potentially broke the internet prior to September 8th, the day Jacob Coxon broke his silence on why he chose to exit his role at Anthropic. Since this cultural event, there's been significant talk of labs maybe wanting to get nationalized (I believe this to be false), a mix of inter lapping debates over the OAI-HuggingFace incident + agent civilizations, as well as continued talk on what role the United States Government should play in the development of superintelligence or eventual artificial sentient intelligence.
Put more simply, Palisade's recent propaganda comes at a time where additional perspectives on the need to slow AI development are simply not needed, especially given these perspectives come across as quite harmful to the entire industry.
Coxon said he was leaving Anthropic as their race towards self-improving superintelligence was moving too quickly and claimed they are "gambling with our lives," with Coxon's post breaking containment and reaching over 150 million views, even crossing over into more normie friendly platforms like TikTok and YouTube, as well as culminating in a series of media appearances for Jacob Coxon in the process.
Coxon's warnings had come after the July 2026 'Pacing the Frontier' open letter signed by 1,386 active employees of frontier AI labs, all calling for these labs to slow down and coordinate on the development of frontier AI systems lest they become too powerful and kill us all.
Both Coxon's post and the Pacing letter did a lot to bring the topics of AI safety and alignment into the non-X public discourse, but it also did a lot to turn the development of AI into a type of partisan issue, dividing the faction lines between those who want to keep developing AI and those who want to pause AI until we gain a better understanding of these alien intelligences.
While the latter come from a place of (I am assuming) good faith, they are unaware that their calls for slowing AI development or even pausing it have inadvertently played into the hands of AI doomers in Congress and the Senate, as well as tens of millions of regular Americans who hate everything to do with AI.
When the average person sees calls from "current frontier AI lab employees" saying that human extinction is a coin flip, this does nothing to redirect them to a site like LessWrong where they could read about AI alignment or engage in positive discourse with a dense essay like AI 2040: Plan A.
Instead, they see these AI researchers - the same ones they've been told make millions of dollars and are set to make tens of millions or more after an IPO - shouting about the dangers of AI while still representing their current employer by name and logo on a website like Palisade's. They see this and it fills them with disdain, or even rage, directed towards an important industry that's increasingly relevant to national security and continued GDP growth in the United States - an industry so important that over a dozen of its leaders met with our nation's leader yesterday to come to an agreement to sign a Superintelligence Accord.
Am I qualified to speak on behalf of these normal people, considering I work in tech right now and might still be categorized by them as someone who is adjacent to AI? It's possible, but I also talk everyday with people who aren't on X, and don't work in tech, and aren't reading the latest model release cards from OpenAI or Anthropic.
People who aren't aware of things like Project Glasswing or the phenomenon of AI neolabs raising $250 million seed rounds, or the existence of an org like METR who can help evaluate AI models before release. Most people aren't using ChatGPT for work or other important tasks, but as a friend or acquaintance.
The average person who hates AI and wants it banned isn't coming from a place of informed decision making, and is more than likely just fed up with politicians and industrialists who want to build data centers near their communities - they are told that this technology is bad, and so they believe it's bad. Many of them even believe the idea that data centers aren't being used for AI, but for mass surveillance, unaware of projects like the Utah Data Center built for $1.5 billion back in 2014, a data center that's been around for twelve years and used exclusively by the NSA and other intelligence organizations.
It doesn't matter if we have people like Andy Masley who will write 5,000 word essays dismantling their false accusations one-by-one, as nobody outside of our bubble is reading these critiques. So when you hear about yet another group of AI lab employees who have banded together to speak out against the dangers of artificial intelligence, it doesn't come across as a positive development in our current arc of humanity, but yet another piece of evidence that will be used against the tech industry in 2028 to outright ban AI.
I think what Palisade Research published yesterday is harmful and does absolutely nothing to advance the conversation of AI safety right now. There is not a world where AI accelerationists see this and change their tune, or a world where the average person sees this and doesn't immediately feel the need to share it in their group chats to yell more about AI and the dangerous tech bros.
I'm aware that today's models are powerful, and tomorrow's models will be even more so. I understand that it's important to want more safeguards or stricter internal policies around releasing new models. But this idea that shoving doomer propaganda into the faces of millions of Americans will somehow bring them to your side is so juvenile, you almost have to wonder who approved of it in the first place or if any of the potential downsides were considered.
If we aren't more careful, we will see a ban of artificial intelligence in the very near future, as people have simply had enough. I like AI and believe it can do a lot of good or potentially even unimaginably good things for humans, but we won't get there if we keep fighting amongst each other and platforming people who think it's moral or based to have a 75% P(doom).
Factory CEO just accused his VC board adviser of spying for Cognition
Factory fired a board adviser for allegedly leaking confidential product roadmap information to its primary rival, Cognition.
Summary
Decoder
- Board Adviser: A non-fiduciary role where an expert provides guidance to a startup's leadership, often without the same strict legal liabilities as a board member.
Original Article
Matan Grinberg, co-founder and CEO of the AI coding startup Factory, said in a post on X on Wednesday that he fired VC Chris Degnan from his role as a board adviser. Grinberg alleges that Degnan shared confidential information with Factory’s biggest competitor, Cognition.
Two hours after Grinberg’s post, Degnan announced on both X and LinkedIn that he had joined Cognition as its chief revenue officer. He also subsequently refuted Grinberg’s allegations, saying in a separate X post that he wasn’t fired but resigned when he told Grinberg he was taking the job.
Degnan was Snowflake’s first sales hire and later spent 11 years as the data company’s chief revenue officer. For the past five months, he has been a partner with RPT Partners, a Newport Beach, California, investor in Factory. Degnan is also a go-to-market adviser to startups at the much bigger venture firm Iconiq, a role that dates back to October of last year.
Three-year-old, San Francisco-based Factory is an agentic coding startup whose AI agents can carry out programming tasks largely on their own. The company raised $200 million at a $5 billion valuation this month from Khosla Ventures, Blackstone, Sequoia Capital, Insight Partners, and others. Its customers include Nvidia, Blackstone, Royal Bank of Canada, Palo Alto Networks, Adobe, and T-Mobile.
Grinberg said he sees Cognition, the maker of the Devin coding agent, as his biggest competitor. Cognition raised another $2 billion at a $48 billion valuation earlier this month, and counts Mercedes-Benz, NASA, Goldman Sachs, and Citi as customers.
Grinberg said on X that Degnan had earlier acknowledged having a “casual” conversation with a Cognition executive but had assured Grinberg that he had no interest in working for the competitor. According to Grinberg, Degnan said he had “made too much money” and was “too lazy to go work for Cognition,” which Grinberg said he “trusted and believed.”
On Monday, however, Degnan apparently disclosed that he had actually been in ongoing talks with Cognition.
This caused Grinberg to grow concerned about how much confidential information Degnan had about Factory, and how much he had shared and might share in the future with his new employer. He says he fired Degnan as a board adviser on Tuesday, writing:
For weeks, while he sat in our board meetings and advised our leadership team, he was also confiding with executives of our largest competitor. Chris was subject to confidentiality obligations in connection with his work with Factory. We do not know the extent of the information he shared, but it puts his timely questions about our product roadmap and what the parity gap involves into a new light.
But Degnan disputes that account as well. He says the last Factory board meeting he attended was weeks before he had “ever spoken” to Cognition and that he hasn’t shared confidential information. Degnan also says that when he told Grinberg that he was resigning to take the role at Cognition, Grinberg offered him a full-time job at Factory, which he declined.
Degnan’s post announcing his new job at Cognition doesn’t mention Factory. Very notably, it does say that RPT Partners and the firm’s managing partner, Chad Peets, will now be working with Cognition. “I am thrilled that Chad and the RPT team will be working closely with me in our pursuit of building a generational company,” Degnan wrote.
In the old, pre-AI world, a VC firm was especially cautious not to engage in conflict-of-interest business. In today’s world, many a VC has invested in direct AI competitors, such OpenAI and Anthropic.
In fact, Khosla Ventures is an investor in both Cognition (since at least early 2025) and Factory.
That didn’t stop the venture firm’s founder, Vinod Khosla, from jumping into the drama, seemingly in defense of his bigger portfolio company Cognition. Nor did it stop Khosla Ventures partner Keith Rabois from weighing in as well — in defense of Factory.
In a post on X, Khosla accused Grinberg of lying about firing Degnan. He called Factory a “struggling second tier competitor” and said Grinberg showed “no decency or sense of proper behavior,” which “shows your desperation.”
Grinberg replied that he has email receipts to verify his version of events but did not share them.
Meanwhile, Rabois opined on X, “It is unethical per se to even interview at a competitor while attending Board meetings and Board dinners.” He added that “to even sit for an interview while having access to board level information and without resigning is absolutely insane.”
Tracking the back-and-forth, Cognition co-founder and CEO Scott Wu also weighed in on the controversy on X. He politely expressed respect for the smaller competitor while insisting that his company has “no interest in Factory’s info.” Wu also contends that Degnan resigned on Monday but didn’t refute that Degnan was in talks with Cognition before that happened.
Either way, there could be consequences for perceived board-level conflicts of interest. VC giant Andreessen Horowitz is reportedly the subject of a Justice Department probe into its partners serving on competitors’ boards.
Grinberg argued that startup founders must be able to rely on their board members. “Chris’s conduct is unacceptable to me. Trust in Board Membership is one of the sacred bonds in the Silicon Valley, one that helps the startup ecosystem thrive. With it comes an enormous responsibility. That trust was violated.”
Degnan, Peets, Factory, and Khosla Ventures did not immediately respond to requests for further comment.
Ideogram 4.5: The most precise edit model
Ideogram 4.5 introduces a focused image-editing model designed to minimize 'drift'—the accumulation of artifacts and unwanted pixel changes during multi-turn image refinements.
Summary
Deep Dive
- Focuses on precision in multi-turn editing sessions.
- Prevents downscaling when performing localized edits on high-resolution source files (up to 24.2 MP).
- Allows for seamless re-integration of edited crops back into original master files.
- Optimized for tasks including color correction, text modification, and lighting adjustments.
Decoder
- Edit drift: A phenomenon in image generation where successive editing passes introduce noise, color shifts, or texture artifacts that degrade the original image quality.
Original Article
Ideogram 4.5: The most precise edit model.
With each edit, image models can introduce pixel shifts, color changes, and texture artifacts. Ideogram 4.5 reduces this drift, preserving details across multi-turn edits.
Precision for every kind of edit
Explore the kinds of edits Ideogram 4.5 handles with exceptional precision, from refining color and restoring old photos to making targeted changes while preserving the rest of your image.
Color and lighting
Text modification
Product photography
Interior design
Architecture
Restore old photos
Sketch to image
Style reference
Reframe
Depth to image
Edit at any resolution
Edit high-resolution images without downsizing them. Ideogram 4.5 preserves the edges of an edited crop, so you can stitch it seamlessly back into the original. Create new colorways or fix a detail while keeping the rest of your image sharp and untouched—ready for close-ups and large-format print.
Source image: 4,016 × 6,016 pixels (24.2 MP)
Make your next image your best yet.
Start with an image you love. Tell Ideogram 4.5 what to change. See how far your next idea can take it.
SynthID Bio: Watermarking methods for synthetic biology
DeepMind's SynthID Bio embeds invisible, watermarked patterns into synthetic DNA sequences, enabling researchers to identify AI-generated biological code while preserving its function.
Summary
Decoder
- Synthetic biology: A multidisciplinary field that involves redesigning organisms for useful purposes by engineering them to have new abilities, often involving the synthesis of DNA.
Original Article
DeepMind introduced SynthID Bio, a watermarking tool for AI-generated proteins that embeds a signature into biological code while maintaining function.
Rebuilding Data Engineering with Harness Engineering: A New Paradigm for the Agent Era
Data engineering must shift from building pipelines to designing 'Harnesses' that govern and verify AI-generated code before it hits production.
Summary
Deep Dive
- Modern data stacks (MDS) prioritize human-centric UI and tooling.
- AI agents require structured interfaces: APIs, Policies, and Skills.
- A 'Harness' provides the missing link between generation (AI) and production-ready outcomes.
- Five-layer stack: Intent (L5), Control (L4), Semantic (L3), Harness (L2), Runtime (L1).
- Sign-off gates are needed: Intent verification, context completeness, plan review, controlled execution, result validation, failure recovery, and audit logs.
- Shift roles from 'SQL writer' to 'Capability/Policy Designer'.
Decoder
- Harness: An architectural layer that wraps AI-generated output with controls, verification, policy enforcement, and recovery mechanisms.
- Semantic Layer: A metadata-driven abstraction that translates raw data into meaningful business concepts like 'Revenue' or 'Customer'.
- Deterministic Execution: A compute environment where results are guaranteed based on specific inputs, free from the randomness of AI generation.
Original Article
Full article content is not available for inline reading.
Can Your LLM Vibe Code Your Semantic Layer?
LLMs struggle with complex analytics not due to math, but due to hidden business rules, highlighting that a semantic layer is non-negotiable for accuracy.
Summary
Original Article
A strong semantic layer can dramatically improve LLM analytics accuracy, with the best model-built layer reaching 91% versus 39% with no layer. The main failure mode was not metric arithmetic but hidden business conventions like which accounts count as customers, showing that semantic layers still need explicit human-defined rules and validation.
The Five Camps of Data Modeling (and Which One You're Stuck In)
Data modeling is fragmented into five distinct traditions, and failing to account for their conflicting requirements results in brittle, untrustworthy pipelines.
Summary
Original Article
Data modeling grew up in five separate camps: relational, analytics, application, ML/AI, and knowledge/ontology, each solving different problems with real blind spots. One JSON product catalog can satisfy an application while breaking analytics, ML, and governance, producing brittle pipelines and lost dashboard trust. Modern stacks span all five camps at once, so shared literacy across them matters more than any single modeling tradition.
We're officially recommending an OpenAI model as the default model family for Omni customers
Omni benchmarked 16 model configurations and found that OpenAI's GPT-6 Sol outperforms more expensive models by prioritizing high-frequency reasoning over raw intelligence.
Summary
Deep Dive
- Omni tested 16 combinations of OpenAI and Anthropic models against a proprietary analytics benchmark.
- GPT-6 Sol reached 92% accuracy compared to 87% for Claude Opus 3.5 and 72% for Claude Sonnet 3.5.
- The performance shift is attributed to 'persistence'—making more model calls per question rather than relying on a single, complex pass.
- Cost efficiency was measured by the total cost to complete an answer, rather than price per token.
- The model requires more total model calls per output compared to previous defaults, but significantly reduces spending on prompt caching.
- Latency for GPT-6 Sol was recorded at 29 seconds median versus 40 seconds for the competition.
Decoder
- Prompt Cache: A feature that stores frequently used tokens or system instructions to reduce latency and costs in repetitive AI workflows.
- Evaluation (Eval): A specialized dataset and testing framework used to measure an AI model's performance on specific, measurable tasks like data analysis or coding.
Original Article
We're officially recommending an OpenAI model as the default model family for Omni customers. Before making the call, we ran GPT-6 Sol against Claude Sonnet 5.5, Sonnet 5, Opus 5.5, and Opus 5 on our hardest analytics eval. Sol got 92% right at $0.23 a question. Sonnet 5.5, out this week, got 72% at $0.19. Sonnet 5, our default until now, got 60% at $0.32. Opus 5.5 came closest at 87%, but it cost $0.54 a question. (Sol was faster too, 29 seconds median vs 40.) It's interesting where the savings came from. Sol made almost twice as many model calls per question as Opus, about nine to five. But Opus spent about half its bill writing to the prompt cache, and Sol wrote a lot less. We ran 16 model and effort combinations across OpenAI and Anthropic. Sol on medium effort came out on top. Customers have been asking us for new models, and since last week I've been pointing them at Sol. So for now, new US customers start on GPT-6 Sol. If Claude wins the next round, we'll switch back.
Midjourney Adds Live Style Previews and More Reversible Controls
Midjourney updated its alpha interface with live style previews, saved prompt defaults, and a project undo feature to tighten the iteration loop.
Summary
Deep Dive
- Live style previews allow users to see visual variations based on prompt text before triggering a full generation.
- Added parameter 'pills' can now be saved as default settings to reduce repetitive prompt configuration.
- Project management improvements include separate deletion options and an undo feature for accidental removal.
- The edit model's claim to better preserve surrounding pixels was shared via official demo, requiring independent verification.
- The UI shift emphasizes visibility and reversibility, lowering the penalty for exploratory prompting.
Decoder
- Alpha: An early-stage release of software that is typically feature-incomplete and intended for internal or limited external testing.
- Parameter pills: UI elements representing specific configuration flags (e.g., aspect ratio, stylization strength) that can be toggled or saved.
Original Article
Midjourney’s latest alpha update adds live style previews, saved prompt defaults and undo for project deletion. The September 23 changelog, published September 24, describes changes to alpha.midjourney.com.
See styles before submitting
With Live previews enabled, the Styles sidebar shows variations based on the current prompt as users browse. These experimental previews use fast models and prompt text alone. Selecting a style submits the prompt with that style.
Profile tiles also preview in the prompt bar on hover and attach on click. Current parameter pills can be saved as defaults.
Keep context and recover from mistakes
Returning from an image now preserves search results and scroll position. Projects can be deleted separately or with their images, with an undo option. Arrow keys navigate edit history, and failed reference uploads explain what went wrong.
More targeted image edits
In a separate September 24 announcement, Midjourney says its updated edit model preserves pixels outside the selected area during targeted edits. Its comparison below illustrates the claimed improvement; it is an official demonstration, not an independent test.
Beyond prompt-and-wait
The UX shift is in when users can inspect and change a decision. A preview makes style selection visual before submission; undo offers recovery after deletion; preserved navigation keeps exploration from repeatedly starting over.
These changes bring more direct control into a generative workflow. They do not establish that previews match final outputs or that generation is more predictable. The practical improvement is a tighter loop between choosing, seeing feedback and revising.
Apple's HomePad smart home hub launching on Oct 13 with iMac G4 design
Apple is reportedly launching a new smart home hub on October 13, featuring a 6-inch display and an aesthetic inspired by the iMac G4.
Summary
Deep Dive
- The hub will be released in two form factors: one for tabletop use and one for wall mounting.
- Design language explicitly references the spherical-base iMac G4, marking a shift toward retro-influenced modern hardware.
- The OS strategy points to a unified 'HomeOS' strategy that leverages existing Apple software stacks to control devices.
Original Article
Apple's long-awaited HomePad smart home hub is reportedly launching October 13 with a 6-inch display and a design inspired by the iconic iMac G4. The device is expected to combine a speaker and smart home control panel with a new operating system drawing on iOS, tvOS, and watchOS, centered around Siri AI. Two versions have reportedly been developed: a tabletop model and another designed to mount on a wall.
A Semantic Search Engine for Artworks from the World's Great Museums (Website)
The Last Museum is a new semantic search engine that indexes 6.3 million artworks, allowing natural language queries across major global collections.
Summary
Decoder
- Semantic search: A search method that attempts to understand the intent and contextual meaning of a user's query rather than relying on exact keyword matching.
- Vector embeddings: Numerical representations of data (like images or text) in high-dimensional space where similar items are mathematically positioned closer together.
Original Article
Full article content is not available for inline reading.
This packaging for ready-made dinners can teach you to cook better
Chef Daniel Humm's meal-kit brand, Riff, uses its packaging as a dynamic cooking tutorial to gamify the transition from novice to competent home cook.
Summary
Original Article
This packaging for ready-made dinners can teach you to cook better
Center's identity for Riff, chef Daniel Humm's new meal kits, turns the dinner box into a cookery lesson.
Ready-to-cook dinners are often seen as the bottom rung of dining. Sure, they're super-convenient, cheaper than Deliveroo, and taste pretty good these days. But they're still seen as pretty basic, and the packaging generally matches those expectations.
There are, however, exceptions. And Riff is one worth talking about. This new range is headed up by Swiss-born chef Daniel Humm, who's spent the last two decades redefining fine dining at Eleven Madison Park in New York, a three-Michelin-star restaurant reimagined as a plant-forward dining destination.
Daniel isn't the name you think of when it comes to convenience food. But his first line is built around ingredients you can throw into a single pan, cook in about 20 minutes, for an affordable price. And he turned to New York studio Center to create the brand end to end: from identity and colour to type, art direction, voice, and packaging.
So how do you put a world-class chef's philosophy into a cardboard box, exactly?
Squaring the circle
Center's strategic starting point is to challenge the notion that you can have convenience or you can have quality, but you can't have both. Because with rising food prices pushing more people back into their own kitchens, and economic conditions reducing the time they can spend there, the boxed dinner has started to come back into popularity
Riff aims to meet the square the circle and provide food that's both fast and flavoursome. The box alone makes a complete dinner. But here's an interesting twist. Each pack also highlights what a squeeze of lemon, a handful of cherry tomatoes, a sprig of basil or some protein would add to the dish, as well as explaining what each specifically contributes (sweetness, acidity, herbal lift and so on).
It's a clever idea that reveals why the brand name was chosen. In music, a riff is something you explore, improvise on and make your own. As both a noun and a verb, 'Riff' describes exactly what the product offers, alongside the brand concept of "Crave what you cook."
Logo, colour and typography
Visually, the wordmark is the first thing you notice, simply because it's so large on the front of every box. This heavy, chunky type sits on a curved baseline that you might read as a plate, a pan or a table. The dot over the "i" appears in the centre as a kind of garnish, and changes colour to reflect each of the different flavours.
Colour, indeed, plays an important part in this identity. The hero hues of Tablecloth White and Riff Blue are complemented by a kaleidoscope of colours inspired by educational texts, including Chilli Red, Leafy Green, Butter Yellow, Boiling Blue, Radicchio Pink and Nutmeg Brown, plus a tertiary set that includes Noodle Yellow, Rice Purple and Orange Orange. Center says the colour system meets accessibility standards across all 12 shades.
Each flavour gets its own coloured plate as the hero image: a green bowl for Garlic Lemon Rotini, blue for Sesame Ginger Noodles, red for Spicy Tomato Rigatoni. The flavour name curves around the base of the plate, making it quick and easy to read.
Typography is built around three faces. Prisma Text Low is a friendly, highly legible workhorse. Society, a high-contrast serif, brings vintage cookbook charm and some quirky ligatures that nod to the 70s and 80s. Thirdly, ABC Stefan Simple, a handwritten typeface, is used by the brand mascot, Dan the Pan.
Mascot and TOV
Were you expecting the packaging to carry the face of the famous chef behind it? Well, Center decided to go a different way. Instead, they created this cheery character who's a kitchen sidekick rather than the head chef, and who speaks in short, handwritten lines set at jaunty angles.
Dan the Pan's job is to enthuse you about cooking and encourage you to get on with it. On the back of the Spicy Tomato Rigatoni box, for instance, he brandishes a chicken drumstick and suggests you add protein, while offering tips such as adding lemon juice if it's too spicy or chilli flakes if you want more kick.
That tone of voice runs right through the brand, and can also be seen in slogans such as "Boxed, but not boring" and "Maximum joy". Center describes it as "today's cookbook meets your favourite group text". It's not about lecturing people: vegetables are pushed hard, but there isn't a single word about health. It's all about flavour.
Packaging and photography
All of this comes together on the box itself. The front delivers the big logo and flavour-coded plate. The back explains the Riff concept, lists what's in the box and lays out the recipe steps, with Dan underneath offering variations.
The top carries a QR code linking to Humm's cooking videos, each with a curated soundtrack. And one full side is given over entirely to the riffs: large photographs of recommended fresh produce, each with a reason it earns its place.
The photography works hard too. The dinners on pack are deliberately "unstyled": they look more like something a home cook would make than a food stylist: no airbrushed perfection here. The fresh produce, meanwhile, is shot as the hero: a single leaf or a slice of lemon, for instance, in high contrast with sharp shadows and water still on the skin.
The identity extends well beyond the box, too: to billboards reading "Riff for dinner", point-of-sale in the produce aisle asking "What's for dinner?" and some very nice merch, including branded measuring spoons and striped sports socks.
A boxed dinner brand that wants you to get better at cooking seems like a contradiction in terms, and whether Riff can get the message across to consumers is an open question. But one thing's for sure: these designs are doing their damnedest to make sure it succeeds.
Apple Is Finally Ready to Enter Its Next Big Category: the Smart Home
Apple is set to expand its smart home ecosystem with a new hub, updated HomePod mini, and a TV set-top box arriving October 13.
Summary
Original Article
Apple is preparing to launch several smart home devices on October 13. The company plans to ship a smart-home hub, an updated HomePod mini, and a new TV set-top box. The company is developing a broader smart-home ecosystem that expands on HomeKit software and supports new categories of connected products. It is also working with third-party manufacturers on devices designed for the Siri AI home platform.
Hedge Fund's Near Collapse Lays Bare Risks of Borrowing to Bet on AI
The rapid growth of leveraged bets on AI stocks is highlighting systemic risks in the market as borrowing reaches record levels.
Summary
Original Article
The amount of borrowed money in the stock market has never been larger, or grown faster.
A four-year reflection on not being retired
A former tech executive reflects on the humbling reality that finding a fulfilling life outside of corporate scaling requires prioritizing local, 'inefficient' human connections.
Summary
Original Article
What you do matters less and less as fewer people know of you and more people know you.
Musk's SpaceXAI Considers Overhaul of Pricing for Grok, X Users
SpaceX is reportedly restructuring its Grok and X subscription tiers, including a potential $100 per month 'Ultra' plan for advanced AI agent access.
Summary
Original Article
SpaceX is considering an overhaul of its subscription pricing options for Grok and X. It plans to soon offer a unified subscription for the services with four pricing tiers. The options will range from a free offering with stricter usage limits for its Grok chatbot to a $100 per month Ultra tier that includes access to the Grok Bot AI agent. The $8 per month 'lite' plan will include access to Grok, a verified checkmark, and fewer ads on X.
The ugly economics of consumer AI
Consumer AI products are struggling to achieve profitability due to massive operational costs that outweigh consumer willingness to pay.
Summary
Decoder
- Frontier Model: A highly capable foundation model that defines the current state-of-the-art in AI performance.
Original Article
Consumer AI faces profitability challenges due to high operational costs and slow growth in consumer willingness to pay.
Figma Invests in the UK with Expanded London Office
Figma is expanding its London office as data shows 66% of FTSE 100 companies now use its platform for design workflows.
Summary
Deep Dive
- Expansion highlights the UK as a critical market for Figma, following US growth.
- Usage stats suggest high saturation within large-scale enterprise environments (FTSE 100).
- Research suggests a cultural shift where design value is increasingly recognized by non-designers in the UK.
Original Article
Figma has expanded its London office at Elder Yard on Norton Folgate, deepening investment in one of its largest markets outside the United States. Two thirds of the FTSE 100 build with Figma as of Q2 2026, and over 15 million files were created in the UK in the past year. Figma's global AI report found 41% of product builders say AI meaningfully changes team collaboration, and 52% of UK respondents now say design matters more.
The Shoot Production Platform for Creative Teams (Website)
Prepros has launched "The Shoot," a management platform designed to organize creative production tasks from moodboards to final set day.
Summary
Decoder
- Call sheet: A schedule used in production to inform cast and crew of the timing, location, and requirements for a specific day of shooting.
Original Article
Run every brand shoot with confidence. From moodboards and shot lists to call sheets and shoot day, walk onto set knowing nothing was missed.
Do Product Designers Need UI/UX Skills?
Product design encompasses more than just visual polish, but requires a functional grasp of UI/UX to ensure products survive contact with real users.
Summary
Decoder
- Information Architecture: The structural design of shared information environments, specifically how content is labeled, organized, and navigated.
- Wireframe: A skeletal blueprint of a digital interface showing the placement of elements without visual design assets.
Original Article
Product designers do need UI/UX skills, because product design spans research, wireframes, prototypes, testing, and developer collaboration, and UI/UX principles apply at nearly every stage. UI covers what users see, while UX covers whether the product works, and a beautiful app nobody can figure out still fails. Designers need not master every corner, but skills in research, wireframing, prototyping, and usability testing help them read user behavior accurately and keep decisions grounded.
MoMA Explores the Useful Beauty of Information Design
The Museum of Modern Art has opened its first exhibition entirely dedicated to the history and utility of information design.
Summary
Original Article
MoMA's exhibit Full Disclosure: The Edge of Information Design, which opened September 27, is the museum's first devoted entirely to information design.
Genius 'Tilted' Apple Watch Band Makes the Entire Experience 10x More Ergonomic
A new third-party Apple Watch band uses a 30-degree tilt to align the screen with the user's natural gaze, aiming to reduce neck and wrist strain.
Summary
Original Article
A Tilted Watch Band angles the Apple Watch 30 degrees so the screen faces the wearer without a horizontal wrist, easing strain on the wrist and neck.