e2e (Website)
The e2e framework reduces AI testing costs by caching agentic steps and replaying them deterministically, combining flexibility with standard testing practices.
Summary
Deep Dive
- e2e allows developers to mix deterministic assertions with LLM-driven actions.
- Agent actions are cached: once an agent successfully completes a sequence of steps, those actions are replayed deterministically in future runs.
- This drastically reduces token consumption for unchanged test paths.
- The framework supports custom AI models via the Vercel AI Gateway, OpenRouter, or direct API keys.
- It provides first-class support for both web and mobile testing (iOS/Android simulators).
- The architecture separates the 'agent' (which decides steps) from the 'engine' (which executes them), allowing for easier debugging and migration from existing frameworks like Playwright or Cypress.
Decoder
- Agentic testing: A testing methodology where an AI model uses its reasoning capabilities to navigate an application UI and interact with elements dynamically, rather than using rigid CSS selectors.
- Deterministic testing: Traditional testing where every step (e.g., click specific button, expect text) is pre-scripted and expected to yield the exact same result every time.
Original Article
Open Source AI testing framework
Customizable TypeScript testing framework for web, mobile apps and more.
import { test, expect } from 'e2e';
test('a member upgrades to Pro', async ({ app, agent, screen }) => {
await app.open('/settings/billing');
await agent.act('upgrade the workspace to the Pro plan');
await agent.assert('the invoice preview shows a prorated amount');
await expect(screen.getByRole('status')).toContainText('Pro');
});
Why e2e
Most AI tests start from scratch. Every single time.
e2e remembers your agent's actions and replays them when possible, without calling the model again.
Less repeated reasoning. Fewer tokens. Lower costs.
Your testing stack. Your call.
Pick every layer. The config writes itself.
import type { E2EConfig } from 'e2e';
import { web } from '@e2e-dev/web';
import { gateway } from 'ai';
export default {
targets: [{ engine: web(), app: { url: '' } }],
agents: {
default: {
model: gateway(''),
system: '',
},
},
} satisfies E2EConfig;
Connect on your terms.
Use your preferred gateway and keep control of how your AI connects.
Your favorite AI? Bring it along.
Choose the model that works for you. Your testing framework shouldn't decide for you.
Your agent. Your rules.
Give your agent a purpose. Define how it should approach testing, what to look for, and what matters to you.
Your app. Your environment.
Test on the web, go mobile, or bring your own engine. One framework, built to adapt.
AI where it helps. Code where it counts.
Don't make AI think twice.
Replay compatible cached actions without unnecessary model calls.
Goals, not scripts.
Describe your testing goals in plain language. Let your agent handle the actions.
Exact when it matters.
Combine AI-driven actions with precise locators and deterministic assertions.
Your AI. Your choice.
Choose your preferred AI model and customize your testing agent.
One API. Multiple environments.
Test web and mobile applications through one shared API.
Bring your own engine.
Extend the framework with your own execution backend.
Questions, answered.
Which AI models can I use?
Any AI SDK model works. Sign in with the ChatGPT, GitHub Copilot, or SuperGrok subscription you already pay for, bring a Vercel AI Gateway or OpenRouter key, or point the runner at a local server. The config selects the model; e2e adds no markup.
Can the model see my secrets?
Secrets stay in environment variables: the config references named credentials and the runner fills them at run time, so values never live in your test code or transcripts.
Does it test mobile apps?
Yes. @e2e-dev/mobile drives iOS simulators and Android emulators through agent-device, with the same tests, fixtures, and assertions you write for the web.
Does it run in CI?
Yes. CI mode switches on automatically: one retry, a read-only trace cache, and agent steps replay verified actions without model calls when nothing changed. Pass your model key as a secret and run it like any other test job.
Can I migrate from another framework?
Yes: the docs include migration guides for Playwright, Cypress, Selenium, Detox, and Maestro. Locators and assertions map closely, and agent steps replace the most brittle parts first.
What happens when a run fails?
You get a report with the exact failing step, a screenshot, a trace, and logs under .e2e. Re-run with --headed or --debug to watch it live; every error code maps to a fix in the docs.
Could this finally be Microsoft's moment?
Microsoft’s new Copilot ecosystem attempt to capture the 'enterprise harness' by centralizing agent management while remaining model-agnostic.
Summary
Deep Dive
- Microsoft's vision for 'the harness' involves centralized identity, plugin management, and enterprise-grade security for AI agents.
- The new Copilot brings together Chat, Code, and persistent agents under one workflow.
- Microsoft is positioning Windows 365 and Managed Runtime as the execution environments for agent-driven software.
- A critical shift is separating the 'AI harness' (the UI/governance layer) from the underlying LLM provider, allowing businesses to swap models as costs or benchmarks evolve.
- Microsoft's September updates include automatic routing between different agent modes and deeper Office 365 integration.
- The company is addressing the persistent worker problem by offering persistent cloud machines for agent tasks.
- Pricing is moving toward usage-based models, but with new administrative controls to manage budgets.
Decoder
- Harness: The infrastructure surrounding an AI model, including interface, memory, tools, identity, and governance, that enables an AI agent to perform persistent business work.
- Agent 365: Microsoft's proposed integration for managing AI agents through existing corporate governance and identity controls.
- MCP (Model Context Protocol): An open standard for connecting AI assistants to data sources and tools.
Original Article
Full article content is not available for inline reading.
Introducing Cloudflare Traces: follow requests through our entire platform
Cloudflare is extending its platform-wide request tracing to include security rules, cache decisions, and origin handling, with full OpenTelemetry support.
Summary
Deep Dive
- Automatically captures request spans including security rule evaluation, cache hits/misses, and origin timing.
- Provides native W3C traceparent propagation to connect Cloudflare spans with application-level spans.
- Implements Trace Rules using Cloudflare's rule language for fine-grained sampling control (e.g., 100% trace on specific IPs or paths).
- Uses the OTLP (OpenTelemetry Protocol) standard for vendor-neutral data export.
- Integrates with coding agents via SQL API for automated production debugging.
- Pricing is based on ingested data volume rather than span count.
Decoder
- OTLP (OpenTelemetry Protocol): A vendor-agnostic data specification used to transmit telemetry data between sources and observability backends.
- Traceparent: A standard header format used to propagate tracing context across distributed services.
- Span: A single unit of work in a trace, representing a specific operation and its timing within a larger request.
Original Article
Today, we’re introducing Cloudflare Traces in open beta, extending automatic tracing beyond Workers to the rest of the request path. In one trace, you can see supported security rules, transformations, cache decisions, routing, Worker execution, and origin handling, then continue that trace through services running on Cloudflare, at your origin, or elsewhere in your stack. This is a long-term investment in OpenTelemetry and in making Cloudflare the most observable part of your stack.
You can now:
- Automatically trace requests across Cloudflare: Capture supported platform operations in one request-level timeline, no additional set up required.
- Control which requests are traced: Set a baseline sampling rate, then use Trace Rules to override it for matching traffic.
- End-to-end trace context propagation: Accept and forward W3C traceparent headers
- Investigate traces in Cloudflare: View request timelines and span details directly in the Cloudflare dashboard.
- Export traces with OpenTelemetry: Send your spans to any destination with a compatible Open Telemetry Protocol (OTLP) endpoint.
You can enable tracing in the Cloudflare dashboard on any domain or let your agent set up for you.
Giving you the visibility we use to debug Cloudflare
When our own teams investigate, we use our own internal traces, which often include thousands of spans for a single trace, generated by dozens of services and features. This lets us dig deep into every detail of a given request. We don’t think that visibility should stop at our internal systems.
Workers Tracing was our first step toward exposing what happens on our platform. Last year, we launched automatic instrumentation for Worker invocations, including outbound fetches and calls to KV, R2, D1, Durable Objects, and other Workers. It shows the work performed inside the Workers runtime without requiring tracing code for every operation.
The goal of Cloudflare Traces is to bring the same level of visibility to everyone using Cloudflare, whether you’re building on Cloudflare or just have Cloudflare in front of an origin. You get to see how your traffic moved through our platform, and connect the dots between how you’ve configured Cloudflare, and how this influences request processing time, routing decisions, and more.
Follow one request end to end
A request’s path through Cloudflare can be complicated! It might pass through security rules, transformations, routing, caching, or proxied to another service entirely. Cloudflare Traces records each supported step as a span, including its timing, outcome, and relevant attributes. Instead of reconstructing the request from separate logs and configuration, you can see the request’s path through our system in one place.
You can answer questions like:
Why was the request blocked or challenged, and which security rule took action?
See when custom or managed rules evaluated the request, how long evaluation took, and the resulting action. Identify the rule responsible for a block or challenge through its span events.
Was the URL rewritten by a Transform Rule before it reached the application?
You can open the http_request_transform span to see each change, the request component it affected, and the rule responsible. You can also see where the transformation occurred relative to routing and origin handling.
Which Page Rules, Snippets, or Workers handled or changed the request?
The workers_routing span shows whether a route matched, which routing type was used, and the matching route pattern.
Was the response served from cache, and where was time spent between Cloudflare, the origin connection, and the application?
You can expand nested cache, upstream, and origin spans to see where the request spent its time. Here, you can see there was a cache miss that went to origin and spent 527ms of the 539ms getting a response.
Configure your tracing
There is no special instrumentation, config, or plugins required. Once tracing is enabled for a domain, Cloudflare generates these spans automatically. This lets you extend the trace through third-party services and back again by adhering to open standards. From there, you can control which requests are traced using a baseline sampling rate and Trace Rules.
Set a baseline sampling rate
You can enable tracing on any domain and set a baseline sampling rate to balance visibility, data volume, and cost. You might trace 1% of requests during normal operation, giving you a continuous view of request behavior without collecting a trace for every request.
Configure Trace Rules
Trace Rules let you keep a low baseline sampling rate while capturing complete traces for a specific investigation. If one customer reports a problem, you can trace 100% of traffic for their hostname, source IP, or identifying request header while leaving everyone else at 1%. Or during an investigation, you could trace 100% of requests carrying a temporary debug header, while leaving all other traffic at the baseline. This lets you reproduce an issue without increasing tracing across the entire domain.
Trace Rules use the same Cloudflare Rules language, so you can target paths, methods, headers, IP addresses, geographies, or combinations of those properties.
Accept and propagate trace context
One of the most common requests we hear is for true distributed tracing: a single trace that follows a request into Cloudflare, through our platform, and onward through the rest of your stack.
Cloudflare Traces can accept a W3C traceparent header from an incoming request, allowing Cloudflare spans to join a trace that began before the request reached our platform. An incoming propagation policy controls whether Cloudflare accepts that context.
Cloudflare can also forward a new traceparent header to your origin. Any other instrumented services can extract that context and continue the trace through APIs, databases, and services running on Cloudflare or elsewhere. To view everything as one connected trace, you can send both Cloudflare and application spans to the same OpenTelemetry-compatible backend.
Export traces to your observability platform
You can export Cloudflare spans over OTLP to a compatible observability platform, where they appear alongside telemetry from the rest of your stack. Configure an account-level destination, then choose which domains send traces to it. This is part of our commitment to OpenTelemetry: Cloudflare represents request activity as OpenTelemetry spans and delivers them using OTLP, keeping the data portable across observability tools.
Let your agent investigate Cloudflare Traces
When you ask a coding agent to debug a production issue, it might inspect your code and run tests, but it may not be able to see what happened to the request in production. With the Cloudflare Observability MCP server, your agent can leverage our SQL API to query your traces (and all of your observability data!), giving it access to your investigation production telemetry.
Let your agent find the right requests, comparing failed traces with successful ones, and identifying where their spans diverge. Since the agent can also inspect your repository, it can connect those findings to the relevant code, narrow down what needs to change, and help put up a fix for you to review.
Pricing
Cloudflare Traces will be a part of the unified Cloudflare Observability pricing model. Instead of charging by the number of spans/events, pricing is based on how much observability data you ingest and how long you retain it. New pricing will take effect across Cloudflare Tracing (and Workers Tracing!) starting December 1, 2026.
| Plan | Included Usage | Retention | Additional Usage |
|---|---|---|---|
| Free | 0.5 GB of ingestion per day | 7 Days | Not available |
| Paid and Enterprise | 50 GB of ingestion 10 GB-month of storage per billing cycle | Up to 1 year (coming soon) | $0.25 per GB ingested $0.10 per GB-month stored |
What's next
Following the open beta, we plan to launch:
- Broader automatic instrumentation: Add more spans across both the HTTP request path (e.g. DDoS rules, Access) and the Workers execution path (e.g. Workflows, Queues, Pipelines).
- Authenticated context propagation: Let trusted callers continue an existing trace without accepting context from every incoming request.
- Ad hoc tracing: Capture a specific request on demand without changing the baseline sampling rate.
- OpenTelemetry API support in Workers: Continue building out our OpenTelemetry APIs to enable adding attributes to existing spans or getting trace context.
- Longer retention: Keep trace data available for up to 365 days for longer-running investigations.
Get started
Follow the Cloudflare Traces documentation to trace your first request and tune sampling with Trace Rules. Cloudflare Traces is available in open beta from the dashboard, through the API, or with Terraform, with support for exporting to an OTLP destination.
Closed-loop incident response: connect AWS DevOps Agent to OpenSearch
AWS now enables closed-loop incident response by connecting AI agents directly to OpenSearch observability data using the Model Context Protocol.
Summary
Deep Dive
- Uses MCP (Model Context Protocol) to expose OpenSearch indices as tools to the AI agent.
- Offers three hosting paths: self-managed ECS Fargate, Bedrock AgentCore, or built-in OpenSearch 3.3+ endpoints.
- Requires a webhook forwarder Lambda to transform SNS alerts into signed payloads for the agent.
- Employs Fine-Grained Access Control (FGAC) to restrict the agent's query capabilities to specific observability indices.
- Automates correlation of logs, traces, and CloudTrail events for root cause analysis.
Decoder
- Model Context Protocol (MCP): An open standard for connecting AI assistants to data and tools in a structured, interface-driven way.
- FGAC: Fine-Grained Access Control, an OpenSearch security feature that allows specific permissions down to the index or document level.
- VPC Lattice: An AWS service that simplifies networking between services across different VPCs and accounts.
Original Article
Full article content is not available for inline reading.
High Availability Is Not Resilience: Why Cloud Systems Fail When It Matters Most
A TLS 1.3 upgrade exposed that high availability is often performative, relying on hidden control-plane dependencies that fail silently.
Summary
Deep Dive
- Control Plane vs Data Plane: Failovers often fail because the orchestration layer (DNS) is not as redundant as the application instances.
- Performative Resilience: Organizations often maintain standby environments they never test, leading to 'rotted' recovery paths.
- Ownership: Resilience requires an explicit 'recovery owner' distinct from the standard on-call rotation.
- Comparison: AWS ARC offers more robust, explicit routing control compared to standard DNS-based failover.
Decoder
- Control Plane: The set of infrastructure responsible for managing, routing, and deciding system behavior, as opposed to the data plane which processes actual user requests.
- RTO (Recovery Time Objective): The target time within which a business process must be restored after an incident.
Original Article
High Availability is Not Resilience: Why Cloud Systems Fail When it Matters Most
Key Takeaways
- High availability can mask a lack of resilience when failures occur outside the architecture's design assumptions.
- Multi-region architectures can still share control-plane dependencies such as DNS health checks, identity and access management, and routing infrastructure.
- Resilience needs explicit ownership, recurring testing, and continuously maintained recovery procedures rather than relying on on-call ownership alone.
- The cost and operational risk of full failover testing can push organizations toward performative resilience, where recovery is assumed rather than demonstrated.
- Resilience in complex distributed systems is probabilistic, so the practical goal is to build confidence by repeatedly testing recovery paths and exposing hidden dependencies.
A team I worked with upgraded their public ingress load balancers to Transport Layer Security (TLS) 1.3 for compliance. Nothing about the rollout indicated a problem: handshakes completed, services remained healthy, and metrics remained flat.
As of this writing, Route 53 HTTPS health checks require the target endpoint to support TLS 1.2. If TLS 1.2 is disabled and only TLS 1.3 is available, the health checker cannot complete the handshake and will mark the endpoint unhealthy. The CDN deemed the region unhealthy and stopped routing traffic to it.
Within the region, everything looked fine. Services were up, load balancers were healthy, applications continued to serve requests, and dashboards showed nothing unusual. The only visible symptom was that traffic had stopped arriving.
Users were transparently rerouted to another region thousands of kilometers away. Latency spiked in affected geographies. The failover region started scaling under load for which it had not been provisioned. Synthetic monitoring eventually revealed a pattern that internal telemetry could not show. The actual issue took roughly forty minutes to isolate. The failure wasn't happening in the application data plane. It was happening in the control plane, which decides where traffic should go.
The system was highly available, but it was not resilient.
This distinction matters more than most cloud architectures acknowledge. In practice, the two terms are often used interchangeably, even in teams with a strong engineering culture. Yet they describe different problems with different prerequisites. High availability is about surviving expected failures with minimal interruption. Resilience, on the other hand, is about recovering from conditions the system was never explicitly designed to handle. Part of the confusion is that availability is measurable in terms of uptime percentages, failover timing, and replication lag. Resilience resists that kind of measurement. Many failure modes emerge only under real pressure. Meaningful recovery tests are often too expensive or too risky to deliberately reproduce.
Availability and Resilience Are Not the Same Problem
Modern cloud platforms make high availability relatively accessible. A managed database with Multiple Availability Zones (Multi-AZ) failover, autoscaling groups behind a load balancer, and a CDN in front helps most systems achieve respectable uptime. Standard patterns work well when failures stay within the assumptions the architecture was built around: an instance crashing, an availability zone disappearing, or a dependency timing out.
The problem is that real incidents rarely fit those assumptions. High-availability engineering assumes failures are isolated and predictable. Resilience engineering starts with the opposite belief. Eventually something critical will fail in a way nobody modelled. Redundancy alone will not protect against failure. A failover path that has never been exercised is not a recovery strategy; it is just an assumption. Three patterns make the gap concrete.
Correlated Failures Across Redundancy
Modern HA does account for some correlated failures, such as Multi-AZ, which clearly models an entire zone going down at once. What it doesn't account for is software-layer correlation, such as a bad config push deployed to all replicas simultaneously, a poisoned cache record served from every read replica, or a dependency upgrade that silently breaks a contract on which the whole system relies. When independence breaks down at that level, redundancy stops working as insurance.
Graceful Degradation That Isn’t
Most systems are assumed to degrade gracefully, but few have ever been tested under realistic loads. When the read replica falls behind under stress, does the cache layer absorb the pressure or amplify it onto the primary? Until exercised, graceful degradation is just a hypothesis, rather than a guarantee.
Recovery Paths That Have Rotted
The system has been running for two years; the rebuild runbook was written before half of the current dependencies existed. IAM policies have drifted and deployment tooling has changed. Recovery is a code path like any other; unused code paths rot.
The Organizational Problems Appear Long Before the Incident
The resilience problem is both architectural and organizational. The architectural part gets more attention. Diagrams are easier to draw than operational ownership is to execute.
I observed one recurring problem. Recovery ownership is often ambiguous even when uptime ownership is not. Most organizations are clear about who owns availability, such as on-call rotations, SLOs, and escalation channels. Few have equivalent structures for resilience. There is no named recovery owner, no recurring testing cadence, and no process for keeping runbooks current as infrastructure evolves. The gap is rarely a deliberate decision; it is the default when nobody claims ownership.
Part of what keeps that gap open is that closing it properly is expensive. Testing the scenarios that actually matter requires dedicated engineering time, operational risk, and infrastructure maintained solely for events that may never occur. The result is a drift toward performative resilience. The architecture diagram looks right, the standby region exists, and the runbook exists. Actual recovery capability is assumed rather than demonstrated, not because engineers are careless, but because proving it rigorously costs real money, carries real risk, and delivers value only at the most adverse moment.
In practice, the first thirty minutes of a major incident are often lost to dependency archaeology. Staff finds that permissions have drifted, runbooks reference tooling retired a year ago, and nobody is sure who still understands the standby environment. None of these problems is exotic. They are normal operational erosion. Recovery paths decay because they are rarely exercised.
The fix is unglamorous and requires explicit recovery ownership, which is separate from on-call. On-call is about response. Recovery ownership is about preparation, which includes the procedure, the tooling, and the responsibility to keep both current. Conflate the two and you get great incident commanders and runbooks that haven't been opened in eighteen months.
Teams that recover well treat failover drills as a recurring engineering commitment. They start small with one service and one AZ shift. Every exercise finds something broken, such as a revoked permission, a stale runbook, or a silent alarm. You find these issues by running the process. The shift is treating recovery as something you maintain, rather than something the architecture diagram implies exists.
Five Questions for Recovery
- What fails first? Identify the control-plane dependencies your recovery path relies on: DNS, identity, routing, certificates, configuration, orchestration, or third-party services.
- What would the team actually observe? For each failure scenario, define the signal that would tell you something is wrong. A green application dashboard is not enough if the failure happens before requests reach the application.
- Can you execute the recovery path? Test the actual mechanism rather than reviewing it on a diagram. If failover depends on a DNS change, routing decision, or control-plane API, exercise that path under realistic conditions.
- Who owns recovery? The on-call engineer may be responsible for responding to an incident, but someone should also own keeping recovery mechanisms, runbooks, dependencies, and tests current.
- What changed after the last test? Record what failed, what was surprising, and what assumptions proved wrong. Resilience improves when these findings become engineering work rather than remaining incident notes.
The Control-Pane Trap Nobody Draws on the Architecture Diagram
Most discussions about cloud failover focus on the data plane, specifically about where traffic goes, which region handles the request, and what happens when an instance dies. The control plane, the system that decides where traffic should flow, gets far less attention. The TLS 1.3 incident above is a direct result of that gap.
The data plane worked perfectly. The control plane broke.
Figure 1. A control-plane failure can redirect traffic away from a healthy region while the data plane continues operating normally.
You couldn't see any of this from inside the service. The service and the load balancer were healthy, but traffic simply wasn't arriving. Synthetic monitoring showed a diverging pattern that internal telemetry could not surface, because it only observed the data plane. The failure was visible only from outside, through monitoring infrastructure that was itself part of the control plane.
This is the structural problem with control-plane dependencies. They are invisible in normal operation. You don't think of CDN health checks as part of your service. You think of them as monitoring. However, monitoring that gates traffic is part of the data path, even though it lives in the control plane. When it breaks, your data plane is fine and your service is unreachable.
The pattern repeats across cloud services. AWS STS is a clear example. The global endpoint at sts.amazonaws.com is "hosted in a single AWS Region, US East (N. Virginia). Like other endpoints, it doesn't provide automatic failover to endpoints in other Regions," as the AWS documentation explicitly states. Many ostensibly global services carry a regional dependency that doesn't appear in standard architecture diagrams. If that region degrades, the global service can fail in ways that bypass regional redundancy entirely.
This points to something broader. Many multi-region architectures are only multi-region in the data plane. Their recovery assumptions still collapse onto a small number of shared control-plane dependencies. Moreover, those dependencies rarely appear on the architecture diagram until something breaks them.
There is a reason for this layering. The data plane is built to keep working when the control plane is impaired. An Amazon EC2 instance keeps running even if the EC2 control plane degrades. An in-flight AWS Lambda invocation will complete even if the Lambda control plane stalls. That is good. Although when you build failover on top of the control plane, as DNS-based failover via Route 53 health checks does, you inherit those failure modes whether you intended to or not. Your failover only works if the system that decides to fail over is itself working.
The question worth asking when designing failover is therefore not just "What happens if a region fails?" but "What happens if the thing that detects regional failure is itself degraded?" That surfaces dependencies that would otherwise stay invisible until a TLS version upgrade or a configuration drift makes them visible the hard way.
ARC Versus DNS-Based Failover: An Honest Comparison
If control-plane dependencies are an accepted risk, the question is what to do about them. The two mechanisms most teams evaluate are AWS Application Recovery Controller (ARC) and traditional DNS-based failover via Route 53 health checks. They sound similar, but they aren't.
DNS-based failover is the well-worn default. Route 53 probes each endpoint, marks it healthy or unhealthy and updates DNS records accordingly. They are simple, widely understood, and cheap. But if your organization genuinely cares about sub-minute recovery objectives, DNS failover quickly becomes uncomfortable. DNS propagation delays (i.e., TTLs, recursive resolvers, and clients that cache aggressively) all become part of your recovery window, whether you planned for them or not. Add the control-plane dependency described above, and the picture gets worse. If Route 53's health check infrastructure degrades, your failover will either not trigger or will trigger incorrectly.
ARC takes a different approach. Instead of DNS, it provides routing controls you flip explicitly when you decide to fail over, with continuous readiness checks so you know whether the failover target can actually accept load before you flip. The routing control plane is a cluster of five regional endpoints designed to remain operable during regional impairment. ARC doesn't remove complexity. Instead, it relocates complexity into a more explicit and dedicated operational model.
- Speed
DNS is bounded by health check interval, propagation delay, and client TTL for minutes in practice, but sometimes for longer. ARC operates in seconds. If your RTO is in seconds, DNS is not an acceptable choice. - Failure modes
DNS failover depends on Route 53's health check infrastructure’s accuracy; the TLS 1.3 incident is exactly that dependency misfiring. ARC's cluster of five regional endpoints indicates a routing control flip doesn't require any single control plane to be healthy. - Operational complexity
DNS is something every engineer already understands. ARC is operationally heavier. The routing controls and readiness checks require dedicated ownership. Runbooks need to stay current. In addition, on-call need to know how to execute a flip under pressure, not discover the runbook is wrong mid-incident. - Cost
DNS is effectively free. ARC charges per-cluster and per-control; such charges are not expensive for a single critical service, but they add up. - Pre-failover confidence
ARC's readiness checks continuously validate that the target can accept traffic. With DNS, you find out at failover time.
Given the benefits of each tool, a tiered approach works well, with ARC for the small set of customer-facing critical paths, where control-plane dependency is unacceptable and DNS for everything else.
Recovery Is a Capability, Not a Property
Recovery capability is not something organizations discover during an incident. By that point, whatever capability exists has already been built or neglected months earlier.
The difficult part of building recoverable systems is not just organizational: Resilience cannot be fully proved. High availability has benchmarks, including uptime, redundancy, failover timing, and replication lag. Resilience is also adversarial. Correlated failures, stale runbooks, control-plane coupling, and human coordination under stress, many of these issues only surface under conditions too expensive or too risky to reproduce reliably. Full regional failover under production load and recovery from a corrupted primary with real users waiting are not scenarios most teams can routinely exercise without accepting significant operational risk. This is the reason that performative resilience is so common, not because teams are careless, but because the bar for genuine proof is genuinely high.
The honest goal is not to guarantee recovery. In sufficiently complex systems, resilience is probabilistic, not provable. The realistic target is to improve confidence by reducing unknowns, rehearsing coordination, narrowing the blast radius, and shortening the gap between failure and detection. Architectural diagrams look resilient because the components are redundant. However, resilience is not determined by how the diagram looks under normal conditions. It is determined by whether people can actually recover a system under pressure, with degraded visibility, incomplete context, and dependencies that aren't behaving as expected.
Amazon Aurora PostgreSQL now supports direct querying of Apache Iceberg and Parquet data in your data lake
Amazon Aurora PostgreSQL now supports direct SQL queries on S3-based Parquet and Iceberg tables, using an embedded DuckDB engine to eliminate ETL needs.
Summary
Deep Dive
- Uses an embedded DuckDB engine to perform high-performance analytical scans on Parquet/Iceberg files.
- Supports joining live, uncommitted operational data with historical data lake records in a single query.
- Implements schema inference to automatically map foreign tables from S3 metadata.
- Data lake metadata can be fetched via Glue Data Catalog or IRC-compatible catalogs.
- Performance can be audited via
aurora_analytics_stat_statements()to track bytes read and cache efficiency.
Decoder
- Predicate pushdown: An optimization that filters data at the source (in this case, S3 files) rather than loading all rows into the database engine first.
- ETL (Extract, Transform, Load): The process of moving data from source systems to a target warehouse; reverse-ETL moves data from a warehouse back to an operational system.
- Iceberg: An open table format for huge analytical datasets that provides ACID transactions and SQL-like capabilities on top of flat data files.
Original Article
Amazon Aurora PostgreSQL now supports direct querying of Apache Iceberg and Parquet data in your data lake
Today, we’re announcing a new capability for Amazon Aurora PostgreSQL that you can use to directly query operational data together with data stored in your data lake in Apache Iceberg and Apache Parquet formats, using your existing PostgreSQL applications and tools. By eliminating the need to extract, transform, and load (ETL) structured data from data lakes into your operational database, you can reduce operational complexity and simplify application development. You can also use Aurora PostgreSQL to query data from data lakes managed in Iceberg REST Catalog (IRC)-compatible catalogs, giving you access to data across a breadth of analytics systems without moving or duplicating it. Whether you’re powering real-time dashboards, enriching transactions with historical context, or building AI agents that reason over both live and archived data, you can now do it all through a single, familiar interface.
Previously, if your application needed to combine recent transactional data in Aurora with historical records stored in Amazon S3, a common approach was to build reverse ETL pipelines that duplicated data, increased infrastructure costs, and required ongoing engineering effort to keep everything synchronized. This challenge only grows as you increasingly embed AI agents into your applications, where it is impractical to predict and pre-replicate every dataset an agent might need.
DuckLabs, the team that maintains the DuckDB project, recently joined Amazon, and this capability is an example of how the efficiency of DuckDB is being integrated into our services. DuckDB is now embedded directly within Aurora PostgreSQL, so you can query live operational data (including uncommitted writes) alongside your data lake in a single query. Query processing stays within Aurora, with no additional network hops and no ETL pipelines that duplicate data. You can query Apache Iceberg tables managed through the AWS Glue Data Catalog, as well as Parquet and Iceberg data stored in Amazon S3 and S3 Tables. You do all of this using familiar PostgreSQL syntax and your existing applications and tools.
We’re excited to bring the speed and simplicity of DuckDB directly into Aurora PostgreSQL, so you and your agents can query and combine operational and Iceberg data using the familiar PostgreSQL applications, tools, and endpoints already in use. By building this capability around DuckDB, future improvements to the open source engine can continue to bring performance and functionality gains to Aurora and other AWS services.
What is new
This capability is supported on two Aurora PostgreSQL major versions: 17 (starting with 17.11) and 18 (starting with 18.6). To use it, you create an Aurora PostgreSQL cluster, attach an IAM role with the AuroraAnalytics feature, and enable the aurora_analytics extension. The IAM role is what gives Aurora access to your data in Amazon S3 and the AWS Glue Data Catalog. You then create foreign tables that point to your Iceberg or Parquet data in the data lake, and query them using familiar PostgreSQL syntax. You can complete this setup through the Amazon RDS console, or with any PostgreSQL client such as psql. The process is well documented in the Aurora PostgreSQL documentation.
You can query data across external IRC-compatible catalogs through AWS Glue Data Catalog federation. You register the external catalog once with Glue, and then create foreign tables for the tables you want to query, the same way you would for any Glue-native table. A single query can then join data stored in Aurora with Iceberg tables registered across multiple catalogs, so applications get a unified view without moving data or replacing your existing catalog investments.
Aurora also applies optimizations such as predicate pushdown and column pruning so that only the relevant data is read. This keeps queries efficient even as the underlying data grows. Frequently accessed data is also cached in your Aurora instance, so subsequent queries against the same data return faster. You can inspect this behavior per query using aurora_analytics_stat_statements(), which reports metrics such as rows scanned, bytes read from Amazon S3, and cache hits.
To see how direct querying works, I connected to my Aurora PostgreSQL database using psql and created the extension:
CREATE EXTENSION aurora_analytics;
For my walkthrough, I set up a simple financial scenario. I have a recent_transactions table in Aurora with the last 7 days of customer transactions, and a Parquet file in Amazon S3 containing 5 years of historical transaction data. To make Aurora aware of the historical data, I created a foreign table pointing at the Parquet file in S3:
CREATE FOREIGN TABLE transaction_history ()
SERVER aurora_analytics_server
OPTIONS (
location 's3://<my-bucket>/finance/transaction_history.parquet',
format 'parquet'
);
Notice the empty parentheses in the CREATE FOREIGN TABLE statement. Aurora automatically reads the schema from the Parquet file metadata, so you do not need to define columns manually. For workloads with many tables, you can skip creating them one at a time: a single IMPORT FOREIGN SCHEMA statement bulk-creates foreign tables for every Iceberg or Parquet table in an AWS Glue Data Catalog database, inferring schemas automatically.
With both tables in place, I ran a single query that combines the recent operational data in Aurora with the historical data in S3:
SELECT merchant, category, amount, transaction_date, 'recent' AS source
FROM recent_transactions
WHERE customer_id = 'C-1001'
UNION ALL
SELECT merchant, category, amount, transaction_date, 'historical' AS source
FROM transaction_history
WHERE customer_id = 'C-1001'
AND transaction_date >= CURRENT_DATE - INTERVAL '5 years'
ORDER BY transaction_date DESC
LIMIT 15;
The result shows both recent and historical transactions in a single result set. The 7 most recent rows come from Aurora, and the rest come directly from the Parquet file in S3. DuckDB handles the analytical scan of the Parquet data under the hood, while Aurora handles the operational data. That single query would have previously required a pipeline to move the historical data into the database first.
If a query pattern needs single-digit-millisecond latency, you can materialize data from the data lake into a native Aurora PostgreSQL table using familiar commands such as CREATE TABLE AS SELECT, INSERT INTO ... SELECT, or MERGE INTO. The materialized table lives in Aurora and is queried like any other PostgreSQL table, giving you a low-latency path for hot data without operating a separate ingestion pipeline. The read queries can run on any Aurora PostgreSQL instance in your cluster, whether the writer or a read replica, so you can offload analytical scans from your operational workload. The materialization commands write data into Aurora, so they run on the writer instance.
Get started today
Direct querying of Apache Iceberg and Parquet data from Amazon Aurora PostgreSQL is available today in all commercial AWS Regions and AWS GovCloud (US) Regions, at no additional charge. You pay only for the incremental Aurora compute the queries consume and Amazon S3 request costs for reading data lake files.
To learn more, visit the Amazon Aurora features page, read the Aurora PostgreSQL documentation, or try it in the Amazon RDS console. We welcome your feedback through AWS re:Post or through your usual AWS Support contacts.
The Inference Gap
An analysis of 43,261 Anthropic model calls reveals that 'frontier' performance depends on an unstable inference regime that often fails to manifest in production.
Summary
Deep Dive
- Analysis spanned 43,261 model calls over six weeks.
- 39.2% of calls performed no sequential thinking.
- Median thinking tokens per turn significantly lower than benchmark-level compute.
- A non-linear performance penalty exists when inference compute falls below frontier levels.
- Measured a 20-50% month-over-month decline in delivered reasoning across the corpus.
- Found no evidence of 'nerfing' but clear evidence of a fluctuating, unstable inference regime.
- Benchmark performance assumes 'xhigh' or 'max' effort, which production users often lack.
Decoder
- Inference regime: The combination of compute, reasoning tokens, and infrastructure resources allocated by a provider to a specific model invocation.
- Thinking token: An internal, uninterrupted sequence of reasoning steps performed by a model before providing a response.
- Token-doubling band: A metric used to describe the exponential increase in compute required to gain marginal performance improvements.
Original Article
Full article content is not available for inline reading.
How to Test MCP Apps for Accessibility, and Why It Matters
MCP Apps render live UI inside AI clients, requiring developers to build custom test harnesses to ensure that accessibility is not ignored by the protocol.
Summary
Deep Dive
- Accessibility gap: The MCP spec lacks accessibility guidelines, placing the burden of inclusive design entirely on the developer.
- Testing strategy: Create a local dev server that acts as a "stub host" to simulate host responses without actual agent involvement.
- Addressable states: Use URL parameters to force the app into specific states (e.g., error, loading, empty, teardown) for automated scanning.
- Document structure: MCP apps run in iframes; always start headings at
h1rather than nesting them under the host's heading structure. - Keyboard navigation: Binding custom keyboard shortcuts can conflict with host-level controls; prioritize remappability to meet WCAG 2.1.4 requirements.
- Responsive styling: Use custom CSS properties for safe-area insets instead of
env()selectors, as the latter resolve against the iframe viewport rather than the device screen.
Decoder
- Model Context Protocol (MCP): An open-source standard enabling AI agents to connect to external tools and data sources, recently expanded to support rendering UI widgets directly in the chat client.
Original Article
MCP Apps are changing how AI agents show information to people. Right now, when your MCP (Model Context Protocol) server returns a result, the agent handling the conversation decides what to present to the user: what to show, how to phrase it, or whether to show anything at all. MCP Apps change that. Instead of leaving those decisions to the agent, you can ship an actual interface that renders inside the client itself: buttons, forms, cards, real UI.
That’s good news, but there’s a catch. If a tool is going to ship real UI instead of plain text, you need to make sure the UI works for everybody. Otherwise, as MCP Apps spread, we will have built an entire new class of interface that locks people out from day one.
We need to step in right now, because the risk here isn’t just a poor experience. Whether the client also describes your result in text is entirely up to the client. The spec says nothing about it, and behavior varies. In our testing, Claude Desktop described results alongside the widget, while other developers have reported clients that suppress that text completely. You don’t control which one your users get.
A description isn’t a substitute anyway. Text can summarize what your app shows, but it can’t replace what your app does. Every button, filter, form field, and sort control exists only inside the widget. If the widget is inaccessible, those capabilities disappear for people who can’t use it, regardless of what the client writes underneath.
This means we need to test for accessibility.
Unfortunately, however, the official MCP Apps testing guide covers functional testing only: no unit testing, no integration frameworks, no accessibility. On top of that, the MCP specification itself doesn’t mention contrast, keyboard, screen reader, or focus. They’re both essentially silent on accessibility. This issue has been raised before, including in this GitHub ticket from February 2026, which asks the group to “validate that the spec provides the hooks and constraints needed for strong accessibility.”
Fortunately, while testing MCP Apps for accessibility is fundamentally different from testing web apps, it can definitely be done, and done well. In this post, I’ll show you how.
Why testing MCP Apps is different
When it comes to testing MCP Apps for accessibility, you’re immediately dealing with two concrete problems that don’t exist in the same way with web apps.
You can’t force a live agent to show you every state your app can be in. Testing normally means producing edge cases on demand: the empty result, the error, the string long enough to break your layout. Against a real client and a real tool call, you would have to find inputs that happen to produce each one, and for some states, no input will do it.
The protocol guarantees at least one of them: your app has to load, complete a handshake, and report that it is initialized before the client can send it any data, so there is always a moment when you are on screen with nothing to show. Canceled and teardown are the same kind of thing, driven by the person or the client rather than by your data. None of these are edge cases you can schedule by picking a better prompt.
Your app renders inside a sandboxed, cross-origin iframe, in a page you don’t control. To test it at all, you need a client that renders your app in a browser. MCP Inspector works well, and our own MCP server team uses it: it renders your app the way a production client does, and it’s straightforward to point browser tooling at.
From there, most testing tools still can’t see into the frame. Axe DevTools can, giving you a choice: scan the host page and get your app’s issues alongside the client’s, or target the frame and get back only your app’s issues.
Both options leave the real constraint in place: you’re a guest. You need to know the frame path; it’s specific to the client you’re testing in, and it breaks when that client changes its markup. The client decides when your app renders and with what data. Nothing about the environment holds still long enough to be a reliable test target.
A local dev server fixes both problems.
How to test
First, some vocabulary. The application your app renders inside, whether that’s Claude Desktop, VS Code, or something else, is what the spec calls the host. Your app and the host talk over a small JSON-RPC protocol, which is where message names like ui/initialize and host-context-changed come from.
An MCP app’s state is driven by the messages the host sends it, so testing one largely means replaying those messages yourself. Apps that call back out, whether through the host or to an API you’ve declared, also need those responses stubbed. You’re not mocking internals, adding test hooks, or building a second version of the UI for testing. Your app can’t tell your harness from a real client, which is exactly what makes the results meaningful.
Setup
Serve the exact artifact your MCP server serves. Import the same HTML from the same module your ui:// resource handler uses. Don’t copy it or build a variant. If your dev server and your MCP server can drift apart, they will, and you’ll end up testing a UI your users never actually receive. This is the single most important decision in the whole setup.
Write a stub host. This takes less than 100 lines of code. It answers ui/initialize, sends tool input and tool result once the app reports it’s initialized, and replies to ui/open-link. That’s the whole thing: no framework, no dependencies, and it works in any language with an HTTP server.
The harness renders no UI, ever. Your app is the top-level document here, so any buttons or state pickers land in the same DOM as your app, and every accessibility scan reports your own dev chrome as your app’s problems. The stub is a script tag with nothing visible.
Make every state addressable. Your stub needs to be told which fixture to replay, and a URL parameter is a simple way to do it: one URL per starting state, selecting a fixture rather than holding your app’s state. Since scanners take URLs, you get a ready list of targets out of it.
Getting to a state isn’t the same as testing it. Once you’re there, interact with the app the way a user does and scan what you reach. Fixtures exist for the states clicking cannot produce: the failed call, the canceled run, the teardown, and the app on screen before any data arrives.
The states you need to test
Walk the message sequence and take every branch: connecting (initialize unanswered), awaiting data, streaming input, success, empty, error, canceled, teardown, and context change.
Then cross that with a data axis: minimal, many items, and pathological data, meaning long unbroken strings, right-to-left text, and missing optional fields. Cross the whole thing again with light and dark themes.
Turn each state into a URL
Print every URL on boot, and serve them as JSON. The output of setup is a list of URLs that any tool can consume, and that a person can paste one at a time. That’s the handoff point between the harness and the testing.
What to run against each URL
Run automated rules plus advanced rules against every target, at a realistic chat-column width rather than a full desktop viewport. Then run Intelligent Guided Tests.
This is where the harness pays off. You don’t need a new tool built specifically for MCP Apps. You use Axe DevTools the same way you already do, against a target you now control completely.
How to avoid common mistakes
Even with a solid harness in place, there are a few mistakes specific to MCP Apps worth knowing about as you build your app’s UI.
Document structure
Start your headings at h1. Your app is its own document inside that iframe; heading levels are scoped to it, and you have no way of knowing what the client’s outline looks like. Trying to slot in underneath it by starting at h2 or h3 is a guess that will be wrong in some clients, or in some position in the conversation.
Don’t skip heading levels.
Give the document a real and sufficiently descriptive title. A screen reader user moving into your app may hear nothing else.
Keep the title and the h1 from saying the same thing. The title should describe the app, something like “Accessibility results,” and the h1 should describe this particular result, like “4 accessibility issues.” If a client announces both, they complement each other instead of repeating. Avoid putting the tool’s name in your h1, because that is the string a client is most likely to be using for the frame label already.
Keyboard shortcuts
If your MCP app binds keyboard shortcuts, be sure to check them against the client’s. If the MCP client has shortcuts to open menus or perform toolbar actions, theirs may trump yours.
Another consideration is that your app lives in its own document, so a keystroke inside it won’t necessarily reach the client through the page the way it would in a typical app. Some clients may forward keys deliberately; others won’t. You can’t assume your bindings work globally, and your users can’t assume the keys they rely on elsewhere in the conversation still work while focus is in your widget. Test both.
Either way, the best approach is to allow users to remap keyboard shortcuts. If you bind a single key, WCAG 2.1.4 requires that the shortcut be switchable off, remappable, or active only while the relevant component has focus. Remapping is the most forgiving option, and it’s the only one that also helps with the collisions above. Whatever you bind, make sure focus can still leave your app by keyboard.
CSS and styling guidance
Use tokens for both halves of a color pair, or neither. The specific failure to watch for is a host-provided background paired with your own hardcoded text; a pairing nobody can verify. The spec itself warns about this hazard.
Pair within a semantic family. Shipping --color-text-danger alongside --color-background-danger says the two are meant to be used together and to meet contrast when they are. --color-text-tertiary on --color-background-secondary says nothing of the kind. The spec doesn’t require hosts to verify either way, so treat a matched pair as a good default rather than a promise, and scan it.
Set color in as few places as possible. Establish a pair on a container, let descendants inherit it, and use currentColor for borders, icons, and SVGs. Fewer declared pairs means fewer combinations you have to verify.
Fallback colors need to work in both directions: legible against your own fallback background, and legible against whatever the host actually provides.
Use --color-ring-* for focus indicators, and check that ring against the host’s surface, not just your own.
Safe area insets
Don’t use env(safe-area-inset-*) directly. It resolves against the viewport of the browsing context where it’s evaluated, and inside your app that is the iframe, not the device screen. Your frame’s viewport carries none of the device’s notch or home indicator geometry, so in practice you get zeros that tell you nothing about where the real edges are. The client is the only party that can see them, which is why it passes them to you instead: read them from hostContext.safeAreaInsets and set them as CSS custom properties yourself, as the snippet below does. The bullets that follow assume you have.
Apply insets to whatever actually paints at the edge, not just body. That includes fixed toolbars and anything else pinned to the edge of the screen.
Compose with max() rather than adding: padding-block-end: max(12px, var(--safe-bottom)). You keep your own padding when the inset is zero and get the inset when it’s larger, instead of stacking one on top of the other.
Reapply insets on every update instead of reading them once. Rotation and display-mode changes both change insets, and host-context-changed sends a partial context, so merge new values in rather than replacing what you have.
Default to zero, and treat the whole thing as progressive enhancement. The field is optional, and plenty of hosts won’t send it.
This mostly matters in fullscreen and picture-in-picture. An inline widget sitting in a scrolling chat column is already inside the host’s own safe area, which is why inline insets are usually zero. The field earns its keep when your app owns the whole screen.
There’s no SDK helper for this. You have to write it yourself:
function applyInsets(insets = {}) {
const root = document.documentElement;
for (const side of ["top", "right", "bottom", "left"]) {
root.style.setProperty(`--safe-${side}`, `${insets[side] ?? 0}px`);
}
}
:root { --safe-top: 0px; --safe-right: 0px; --safe-bottom: 0px; --safe-left: 0px; }
.page { padding-inline: max(16px, var(--safe-left)) max(16px, var(--safe-right)); }
.toolbar { padding-block-end: max(12px, var(--safe-bottom)); }
This mirrors how you’d already write env() in a normal web app.
Reduced motion and other media queries
Reduced motion needs no special handling. Use the same prefers-reduced-motion media query you’d use in any other web app. The same is true for prefers-color-scheme, forced-colors, and prefers-contrast. All four propagate through intact.
What you can do today
MCP Apps are still new enough that whatever approaches we can establish now will be the ones that stick over the long term. Which means if we can get accessibility right at this stage, it will stay right as the ecosystem grows. This is how we can avoid introducing the same problems we’ve had to deal with across the rest of the web for so long.
The good news is, you can start right now. Stand up the dev server. Walk your states. Point the accessibility tools you already have at a target that will hold still long enough to test.
Remember, the ultimate goal is digital products, services, and experiences that work for everyone. MCP Apps are showing promise, and they have a lot of momentum. Now is the time to ensure they’re accessible.
Prime Inference: Fast, Reliable Serving for Frontier Open Models
Prime Intellect is launching a serverless inference platform designed to handle large-scale agentic workloads by disaggregating prefill and decode tasks.
Summary
Deep Dive
- Prime Inference supports both serverless and reserved capacity, using NVIDIA Blackwell and InfiniBand-connected clusters.
- The architecture utilizes NVIDIA Dynamo for orchestration and vLLM for model execution.
- They implemented 'chunked prefill' to prevent long-context prompts from stalling token generation for ongoing sessions.
- KV-aware routing keeps session history localized to specific decoders to maximize cache reuse.
- They contributed a native NVFP4-sparse MLA kernel to FlashInfer, increasing KV-cache capacity by approximately 50%.
- The BLHNC block-major layout significantly reduced NVLink transfer descriptor count and latency compared to layer-major layouts.
- Structural-tag support was added to Dynamo to enforce tool-call schemas and reduce error rates in agentic workflows.
Decoder
- KV Cache: The cached attention keys and values for previously generated tokens, essential for avoiding recomputation in multi-turn conversations.
- Disaggregation: Separating the computational workloads of prefill (prompt processing) and decode (token generation) to optimize hardware utilization.
- Tensor Parallelism (TP): Distributing model weights across multiple GPUs to reduce per-GPU memory usage and latency.
- MLA (Multi-Head Latent Attention): An attention architecture that compresses KV caches to reduce memory footprint at the cost of slight precision loss.
- NVFP4: A 4-bit floating point format used for compressing KV cache entries.
Original Article
Prime Inference: Fast, Reliable Serving for Frontier Open Models
Prime's mission is to build frontier open models and the open superintelligence stack for continuously improving agents. We already provide end-to-end post-training infrastructure, from prime-rl and verifiers to sandboxes and RL environments. But the continual learning loop is not complete until a trained model can serve real users, generate new experience, and feed those production traces back into training.
We're excited to release Prime Inference today. It covers both serverless endpoints and reserved capacity and offers resilient serving of frontier open-source models on our GPU infrastructure across multiple datacenters.
Prime Inference began as the serving platform we needed ourselves. Long before public release, it powered large-scale RL rollouts, synthetic data generation, evaluations, and long-running coding agents, processing nearly a trillion tokens every day just internally. Beyond our own workloads, we've also been serving large-scale customer deployments in production since January. This scale pushed us to optimize for sustained performance, quality, and reliability, rather than benchmark speed alone.
Our first public deployment, GLM-5.3, went live on OpenRouter on September 22. It currently ranks among the fastest GLM-5.3 endpoints on OpenRouter, with a near-zero tool-call error rate and 100% uptime since launch.
Prime Inference at a glance
- Low-latency: Our GLM-5.3 endpoint on OpenRouter is continuously evaluated for quality, with production SLAs, security, and privacy built in from the start.
- Premium infrastructure across data centers: Prime-hosted models run on NVIDIA Blackwell today, with Vera Rubin coming soon.
- Uptime: Automatic failover across data centers keeps traffic moving to healthy deployments.
- OpenAI compatible: Connect your existing tools and SDKs using a Prime endpoint and API key.
- Scaling: Serverless endpoints for variable demand, with reserved capacity for sustained workloads.
- Cost: Unified billing and team-level usage tracking across models, making inference spend easier to manage.
- Robust open-source infrastructure: Our stack combines NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer, developed in close partnership with Inferact and NVIDIA, with improvements contributed upstream.
Get started
prime inference chat 'z-ai/glm-5.3' "Write a haiku about KV caches."
Or point any OpenAI SDK at https://api.pinference.ai/api/v1. See the docs for the full API reference.
Built for production SLAs
Prime Inference separates the public API from the model fleet, so capacity can move, fail, or scale without changing the client endpoint.
Our shared circuit breakers let every gateway replica react to failures consistently, while lease-based admission control prevents overload and automatically recovers capacity when a process disappears.
Below the software layer, every cluster is continuously monitored, with health checks that reach all the way down to NVLink and InfiniBand. Alerts go to an on-call team staffed 24/7, so GPU failures are caught and repaired quickly instead of slowly degrading service. And because Prime maintains significant overflow capacity, we can reroute traffic and bring up new deployments whenever more capacity is needed.
How we serve production agent traffic
Workload
A typical agent turn adds about 6K tokens to a 140K-token prompt, reusing most of the conversation history. Under load, these returning sessions run alongside new requests with long, uncached prompts.
We benchmark this mix with AgentX from SemiAnalysis, which replays multi-turn agent sessions. Our benchmark harness also injects cold arrivals with long prompts. We measure end-to-end tokens per second per user for interactivity and output tokens per second per GPU for efficiency.
Prefill/decode disaggregation
On shared GPUs, processing a long prompt can interrupt token generation for existing sessions. Chunked prefill limits these interruptions, but both workloads still compete for GPU time.
We run prefill and decode on separate GPU groups. NVIDIA Dynamo handles routing and orchestration, while vLLM runs the model on each group. Once prefill finishes, the decoder pulls the computed KV through NIXL and adds the request to its batch.
With Dynamo coordinating separate prefill and decode pools, we reduced p90 inter-token latency by nearly 40% in our tests.
Caching and routing
Dynamo's KV-aware router chooses a prefill worker based on how much of the prompt it already has cached and how much work is queued there. Workers publish cache updates so the router can track where prefixes are available. We also keep sessions on the same decoder between turns to support KV reuse.
Mooncake provides a second cache tier in host DRAM. Prefixes offloaded from GPU memory can be retrieved instead of recomputed, allowing us to retain more conversation history.
Performance: GLM-5.3 on GB200 NVL72
Long-context agentic serving is as much a cache-management problem as a compute problem. Performance therefore depends on retaining that history, scheduling new work promptly, and moving cached state without interrupting ongoing generation.
Our interactivity target was 100 end-to-end tokens per second per user. We tune for the number of concurrent sessions we can support at that speed.
We optimized these paths separately: prefill topology and scheduling to reduce time to first token; compressed KV and a fused attention kernel to support low-latency decoding; and a transfer-friendly cache layout to reduce the overhead of moving KV between workers.
Technical deep dive
The sections below cover our work on topology, scheduling, compressed attention, and KV transfer:
- Choosing the right topology for prefill and decode
- Reducing the scheduler bubble on prefill
- NVFP4 KV compression on FlashInfer
- Faster NIXL transfers on NVLink with the BLHNC layout
Prefill: time to first token
Choosing the right topology
We chose DEP8: eight data-parallel attention ranks with expert parallelism across the group. DEP8 distributes requests across eight attention workers while sharing the model's experts across the group. With the MLA cache layout we used, TEP8 replicated each request's KV across all eight ranks, while DEP8 let the ranks cache different requests. Even after accounting for the extra weight memory this requires, we had roughly five times more usable prefix-cache capacity on the same hardware compared to a topology like TEP8, which was also benchmarked.
Scheduler bubble and the token budget
Cache capacity does not eliminate scheduling delays. We found that a request's cached KV could already be available while the request still waited to enter the running batch.
Workers check for completed loads between forward passes. Results then pass through a batch queue and scheduler, where a ready request can miss the current decision and wait another step. Under load, requests already in progress can fill the next prefill batch, delaying admission even when the cached history is ready. This creates a scheduler bubble.
To reduce this bubble, we halved the number of prompt tokens processed in each prefill step, from 8K to 4K per GPU. Shorter steps let waiting requests start sooner. On our configuration, median queue wait fell from 550ms to 110ms, reducing median time to first token by roughly 20%.
Decode: NVFP4 KV compression
Making room for low-latency decode
We compressed the 512-value MLA latent using NVFP4: four bits per value, with an FP8 scale for each group of 16 values. The 64-value positional component remained FP8. Including scales, each MLA cache row shrank from 576 to 352 bytes.
With the indexer and other state unchanged, total cache capacity increased by roughly 50%, from 1.09 million to 1.63 million cached tokens per decoder at the same memory budget.
A native NVFP4 sparse-MLA decode kernel
We built a native sparse-MLA kernel to remove that intermediate buffer. It consumes the positions selected by the sparse indexer, loads the compressed rows, and unpacks them on-chip as attention needs them. NVFP4 is the storage format; the attention computation uses FP16 operands with FP32 accumulation.
For each query token, a cluster of cooperating thread blocks divides the selected rows. In the GB200 configuration, each cluster spans three to eight SMs. The blocks combine their partial results through distributed shared memory, without a second kernel launch.
NVFP4 KV accuracy
We ran a rigorous evaluation suite, with particular emphasis on long-context tasks, to ensure that NVFP4 KV compression does not degrade accuracy.
KV transfer: reducing copy overhead on NVLink
We use vLLM's NIXL connector to let each decoder pull KV from the prefill workers over multi-node NVLink. In an early comparison, the NVLink configuration added roughly 292 ms to time to first token relative to InfiniBand. The faster interconnect was not translating into faster serving.
The bottleneck was transfer fragmentation.
Changing the layout to reduce fragmentation
We tested BLHNC, a block-major layout recently introduced by the vLLM community. It places a block's data from multiple layers together in memory, allowing NIXL to transfer those contiguous regions with fewer, larger copy operations.
Reliable tool calls
Performance is only part of production agent serving. Agents also need reliable tool calls to read files, run commands, and make edits. If the model calls the wrong tool or produces unusable arguments, the agent must retry or stop.
We contributed a structural-tag builder to Dynamo to translate tool definitions into rules that restrict the names and arguments the model can generate.
vLLM uses xgrammar to enforce those rules during decoding, masking tokens that would violate the tool-call grammar. We also fixed errors introduced during parsing: literal strings such as < were being converted into <, altering code or file contents, while some schema references and nullable argument types were interpreted incorrectly.
Built on open source, with great partners
- vLLM and Inferact. Our stack runs on the vLLM inference engine. We thank Inferact for their deep collaboration on performance and production serving, and vLLM contributors around the world for continually advancing the engine.
- NVIDIA Dynamo. The operating system beneath our LLM inference platform coordinates disaggregated workers, routes requests and KV state, and turns a distributed GPU fleet into one resilient serving system.
On the roadmap
- Batch and async inference for large offline jobs at lower prices.
- Dedicated and 1-click deployments on your own reserved capacity, including your fine-tuned models from Prime training runs.
Aleph Alpha releases open-weight Kolibri with 1M context
Aleph Alpha has released Kolibri, a 78.1 billion parameter Mixture-of-Experts model featuring a one-million-token context window and bilingual English-German support.
Summary
Decoder
- Mixture-of-Experts (MoE): A neural network architecture where only a subset of the model's parameters (experts) is active for any given input, improving compute efficiency.
- Sovereign AI: The capability of a nation or organization to own, control, and maintain the underlying AI models and infrastructure independently of foreign providers.
Original Article
Aleph Alpha has released Kolibri, an open-weight bilingual model aimed at sovereign, mission-critical work in government and regulated industries. The English-German Mixture-of-Experts Transformer carries 78.1 billion parameters while activating 3.46 billion per token, supports contexts up to one million tokens, and can be downloaded with its full weights from Hugging Face under the Apache 2.0 license. Customers can run it on-premises rather than send internal data to a third-party inference service.
Kolibri builds on the earlier Kolibri Origin and Aleph Alpha's automated Model Factory. The new model was trained on 768 B200 GPUs, starting with 20 trillion tokens across 21 days, followed by mid-training and long-context adaptation for nearly 24 trillion tokens in total. Its architecture uses 384 experts with six active per token, with full attention in 10 of 50 layers and a 512-token sliding window in the other 40 to contain inference costs. Although adaptation reached 256,000 tokens, Aleph Alpha provides settings to serve the model at 1,048,576 tokens.
German accounts for 21.3% of pre-training tokens, backed by a bilingual 128,000-entry vocabulary and a tokenizer designed to preserve German compounds. The model offers four reasoning settings, none, low, medium, and high, plus tool calling. Aleph Alpha says it was specialized for German, math, coding, long-context work, and agentic tasks, with sector-specific evaluation suites for public administration, automotive, semiconductors, industrial technology, and aerospace that do not use customer data.
In the company's benchmarks, Kolibri posted an English overall score of 75.5 and a German overall score of 70.8. It scored 96.9 on AIME 2025, 85.9 on LiveCodeBench v6, and 61.4 on BFCL v4 overall. Aleph Alpha says the model sits on the quality-versus-serving-cost Pareto frontier in English and German and can match models with up to four times as many active parameters on math, code, grounding, agentic, and long-context tasks. These are vendor-run results, using Aleph Alpha's own harnesses and the highest available reasoning setting where applicable.
Grounding is central to the release. Kolibri was trained on abstention examples and Aleph Alpha's Merlin-Arthur procedure, which teaches it to withhold an answer when evidence is absent. It avoided a wrong answer on 44% of AA-Omniscience items, compared with 14.8% for Kolibri Origin, and reached 0.23 on the company's M/A grounding score. Aleph Alpha built the model in Germany, trained it in Germany and Finland, and says its control of data curation, training, evaluation, weights, and deployment is intended to meet European compliance and sovereignty requirements. Deployment uses Aleph Alpha's inference package and a Kolibri-specific vLLM plugin.
Sources and related context
- Kolibri deployment with Aleph Alpha’s vLLM plugin: Supports: The official repository documents the inference package and Kolibri-specific reasoning and tool-call parsers used to deploy the model.
- How Aleph Alpha’s Merlin-Arthur grounding method works: Related context: This research explanation describes the adversarial training procedure behind the article’s discussion of document grounding and abstention.
How many AI agents could run on the AI chips shipped through 2027?
Global AI chip shipments through 2027 could support hundreds of millions of concurrent agents, potentially creating an enormous supply of AI work capacity.
Summary
Deep Dive
- The study uses HBM supply as a proxy for inference capacity, calculating effective 'GB300-equivalent' units.
- Concurrency capacity is derived from serving benchmarks (e.g., AgentX) and agent-hour API costs (e.g., Claude Code).
- They estimate a $5/hour rental cost per GB300-equivalent unit.
- Capacity estimates scale with 'u', the performance uplift of newer HBM4/4E memory versus HBM3E.
- Even at 20% utilization, the implied spending levels indicate a massive gap between potential supply and current market revenue trends.
Decoder
- HBM (High-Bandwidth Memory): A specialized, high-speed RAM architecture used in AI accelerators that is currently the primary bottleneck in hardware manufacturing.
- Agentic workload: AI tasks that involve multi-step reasoning, tool usage, and persistent state management, rather than simple token generation.
- Jevons' paradox: The phenomenon where increased efficiency in resource usage leads to increased overall consumption of that resource.
Original Article
Full article content is not available for inline reading.
AI21 uses Kueue to manage a 10,000-GPU fleet
AI21 Labs reports an 83% faster start time for high-priority GPU workloads after migrating to the Kubernetes-native scheduler, Kueue.
Summary
Decoder
- Kueue: A Kubernetes-native job queueing system that manages resource quotas and scheduling for batch workloads, often used to prevent resource contention among teams.
Original Article
AI21 describes replacing ad hoc GPU-capacity negotiations with queueing and fair scheduling across a shared Google Cloud cluster. A Google Cloud case study says high-priority workloads began 83% sooner under the new setup.
Whistle: Speech to Text in 16.9 MB
Whistle is a 16.9 MB speech recognition model that runs locally on CPUs across diverse hardware without external dependencies.
Summary
Deep Dive
- 16.9 MB model size allows execution on low-power hardware.
- Supports seven languages with integrated language detection.
- Uses 80-bin log-mel features with a convolutional stem for initial downsampling.
- Encoder utilizes mHC residual lanes and Monarch Hadamard MLPs.
- Decoder uses a gated cross-attention mechanism to read encoder memory once per clip.
- Decoding supports keyword biasing via Aho-Corasick automaton.
- Benchmarks show superior speed and lower word error rates compared to Whisper base and Moonshine tiny v2 on various datasets.
- Engine supports 17 target platforms including ARM, RISC-V, and WebAssembly.
Decoder
- Aho-Corasick: A string-searching algorithm that locates all occurrences of a set of keywords within a text in linear time.
- Monarch Hadamard MLP: A type of matrix multiplication layer designed to reduce parameter counts while maintaining performance.
- Log-mel: A representation of audio that maps frequencies to a scale that mimics human hearing.
Original Article
Today we release Whistle, a speech recognition model for mobiles, wearables, robots, smart home, automotive and microcontrollers. It is one 16.9 MB file, runs on the CPU with no dependencies, and loads into the same C++ engine as Needle, from the same container and the same quantisation.
Whistle does three jobs, all of them on the device:
- Transcription. 16 kHz mono audio, up to 30 seconds in one pass, in English, German, French, Spanish, Italian, Dutch and Polish. The language is detected unless you name it.
- Word timestamps. Every word with its start, end and probability, aligned from the decoder's attention.
- Speech embedding. The encoder output, one row per 80 ms frame, without decoding a transcript.
The model
The front end. 16 kHz mono audio is framed at a 25 ms window and a 10 ms hop into 80 log-mel bins, band-limited to 250-3500 Hz and normalised per channel. Thirty seconds is 3,000 frames. A convolutional stem of 128 channels and kernel 9 halves that count three times, leaving 375 frames at one per 80 ms. Every stage after this runs at that rate, and embed returns one row per frame.
The encoder. Eight Simple Attention blocks: four mHC residual lanes and a Monarch Hadamard MLP in place of the feed-forward network, the same blocks Needle uses. The attention is not causal. A frame at 3 s attends to a frame at 12 s.
The decoder. Eight Laddered Simple Attention blocks at width 512, 8 query heads to 2 KV heads, 48-dimensional queries and keys, 64-dimensional values, a 3-tap causal convolution on Q, K and V, and engram lookups at layers 3 and 7 over 18,432 slots. That is Needle's block list with a different layer count.
The speech-specific part is one addition per layer. Each decoder layer reads the encoder through a gated cross attention, x ← x + σ(g) · softmax(q̂ K̂ᵀ/√d) V, with a gate learned per layer and K and V taken from the clip. Those projections run once when the clip arrives, 375 frames across 8 layers, and are then held for the whole decode. Five beams therefore cost five short transcript caches, not five passes over the audio.
Decoding. Five beams scored by length-normalised log probability. Keyword biasing walks an Aho-Corasick automaton over the phrases you pass in, alongside the beams, and lifts their log probability as the automaton advances. The transcript is capped at 320 tokens. The vocabulary is 8,192 text pieces plus seven language tokens, one per language, so the detected language is emitted as a token rather than returned out of band.
The ladder is on the decoder. Every depth from 2 layers up was trained as a model of its own, and --audio-depth selects one at load time. The encoder is never sliced: all eight blocks run at every depth.
Silence. The engine measures the clip's loudness range before the decoder starts. Below the threshold it returns an empty transcript and an empty language, and never enters the beam search.
Benchmarks
Whistle is ahead on LibriSpeech test-clean and test-other, on SPGISpeech, on Earnings-22 and on the FLEURS average. Whisper base is ahead on TED-LIUM, on AMI and on the MLS average, at 145.3 MB against 16.9.
Each model ran on its official runtime at its defaults: Whistle's C++ engine at 5 beams, openai-whisper, and moonshine-voice non-streaming over whole audio. Time to first token is audio in to first token. Decode is tokens divided by the wall time after it, so the encoder is not counted twice. Whisper pads every input to 30 seconds, so its time to first token is flat across clip lengths. Whistle's tracks the clip: 5.9 ms at 5 seconds, 11.1 ms at 10, 36.3 ms at 30.
Word error rates are scored with the Whisper normalizers. Whistle's are measured over 86,174 utterances. Whisper's and Moonshine's are the figures their authors published, from the multilingual checkpoints rather than the English-only ones. No test audio appears in Whistle's training or validation data, verified by comparing audio checksums and speaker IDs across every reported test set.
One engine, three ways to load it
needle_load reads whichever model a .cact file holds, so the same binary does speech, text, or both:
needle --model whistle.cact --audio clip.wav
needle --model needle3.cact --tools tools.json --prompt "turn off the kitchen lights"
needle --model needle3.cact --model whistle.cact --tools tools.json --audio clip.wav
On the third line needle_complete takes the clip directly. The engine transcribes it, answers the transcript against your tools, and returns one JSON object with the calls and the speech fields, the speech ones prefixed audio_. No transcript is handled by the caller.
{"function_calls":[{"name":"set_lights","arguments":{"room":"kitchen","on":false}}],
"confidence":0.94,
"audio_text":"turn off the kitchen lights",
"audio_language":"en"}
Get started
pip install cactus-needle
import needle
print(needle.transcribe("clip.wav")["text"])
# turn off the kitchen lights
A 16 kHz WAV or raw samples need nothing beyond the base install. Other sample rates and microphone capture need the [mic] extra, which adds soxr and sounddevice.
Every call returns the text, the language, the milliseconds to the first token and the decoder's tokens per second after it. word_timestamps=True adds each word with its times and probability. keywords=["Siobhan", "Krzysztof"] raises the log probability of those phrases during the search. language="de" forces the language instead of detecting it. needle.Whistle() is the same model as an object, for embed(audio) or to hold one tuned .cact.
needle whistle playground transcribes from the microphone in the terminal, and needle whistle compare runs the same clip through Whistle, Whisper and Moonshine side by side with their timings.
Deploy
The engine ships prebuilt for seventeen targets, from macOS and Linux through Android, iOS, watchOS, Windows on ARM, RISC-V, MIPS, the browser and a WASI component. Every folder holds a needle binary, libneedle.a and needle.h, and loads any .cact you hand it.
needle download macos-arm64
needle download whistle
./macos-arm64/needle --model whistle.cact --audio clip.wav --audio-word-timestamps
needle_load, needle_transcribe and needle_embed are the whole speech C API. The engine reads no environment variables. Every behaviour is a compiled default or an explicit flag.
Weights are on Hugging Face, the engine and its platform folders are in Cactus-Compute/needle3, and the source is on GitHub.
Vx (Website)
Vx is a systems programming language that encodes hardware topology and memory placement directly into its type system to prevent low-level runtime errors.
Summary
Deep Dive
- Integrates hardware topology into the type system to catch bugs like invalid device pointer dereferences.
- Uses linear types to enforce ownership and prevent use-after-move errors on buffers.
- Compiler reads machine files to validate capacity and memory hierarchy limits before binary generation.
- Emits MLIR (Multi-Level Intermediate Representation) for downstream compilation to native code, PTX, or CoreML.
- Designed for static, data-oriented codebases rather than dynamic frameworks like PyTorch.
- Distinguishes between SI and IEC units to prevent common unit conversion errors.
Decoder
- Heterogeneous computing: Systems that use multiple types of processors, such as CPUs combined with GPUs or NPUs.
- SMT prover: Satisfiability Modulo Theories, a class of automated reasoning tools that verify mathematical statements about code.
- MLIR: An intermediate representation infrastructure used by compilers to bridge the gap between high-level source languages and hardware-specific machine code.
Original Article
One Language, Every Chip
Vx is a systems programming language for heterogeneous computing. CPU, GPU, NPU and accelerator memory are part of the type system — so a host thread dereferencing a device pointer is a compile error, not a segfault at three in the morning.
curl -fsSL https://vxlang.org/install.sh | sh
macOS on Apple Silicon and Linux x86_64.
Heterogeneity belongs in the type system, not in the runtime.
Where data lives is part of its type
Most languages treat the accelerator as infrastructure: you write math, and a large opaque runtime decides how to ship it. Vx treats it as semantics. A tensor pinned to NPU high-bandwidth memory has a different type from one in host DRAM, and crossing between them takes an explicit transfer() — even when the hardware boundary is free.
On Apple's unified memory that transfer compiles to almost nothing. It is still written down, because data locality should be provable by reading the source rather than by profiling the binary.
// Two matrices already resident in NPU memory.
fn custom_matmul(
a: Pinned<Tensor<f32, [4, 4]>, Topology::NPU[0]>,
b: Pinned<Tensor<f32, [4, 4]>, Topology::NPU[0]>)
-> Verified<Tensor<f32, [4, 4], Memory::NPU_HBM>> {
let mut result =
Tensor<f32, [4, 4], Memory::NPU_HBM>::uninit();
// Dispatch the computation to the accelerator.
spawn on(Topology::NPU[0]) {
for i in 0..4 {
for j in 0..4 {
result[i][j] = 0.0;
for k in 0..4 {
result[i][j] += a[i][k] * b[k][j];
}
}
}
}
return Verified(result);
}
What the compiler rules out
Vx front-loads into type checking a class of bug that normally surfaces as a runtime crash, silent corruption, or an out-of-memory at training step 1200.
Address-space typing
Dereferencing a device pointer from the host. A Pinned<T, NPU_SRAM> escaping into a host expression.
Capacity admission
A placement whose working set cannot fit the memory space it targets — checked against the declared machine, before a binary exists.
Seam contracts
Reading a buffer whose asynchronous transfer has not been made visible. Discharged by an SMT prover.
Linear types
Use-after-move of a consumed buffer, alongside a borrow checker with variance and region tracking.
Topology reachability
A transfer between two memory spaces with no declared path between them.
Autodiff
Differentiating through a region whose adjoint is not defined.
The machine is declared, not assumed
Most compilers hard-code a cost model. Vx reads one. A machine file describes the memory hierarchy and interconnect of a real part, and the compiler admits or rejects placements against it.
Units are exact integer conversions, never floats: SI prefixes are decimal (GB = 109), IEC are binary (GiB = 230). A figure copied off a vendor sheet means what the sheet meant.
The repository ships machine files for H100, H200, B200, A100, MI300X, Apple M4 and multi-GPU nodes — each citing its sources, and marking unverified figures as unverified.
Memory HBM { capacity: 80 GiB, bandwidth: 3.35 TB/s,
managed: explicit, scope: device }
Memory L2 { within: Memory::HBM, capacity: 50 MiB,
bandwidth: 12 TB/s, managed: cached }
Memory SMEM { within: Memory::L2, capacity: 228 KiB,
bandwidth: 128 B/cyc, clock: 1.98 GHz,
replicas: 132, granule: 1 KiB, scope: sm }
Topology Device {
arch: nvptx64,
memory: Memory::HBM,
transfer Memory::CPU_DRAM -> Memory::HBM : 63 GB/s,
}
How it compiles
A data-oriented parallel frontend
Every symbol, nominal type and monomorphized variant is a flat 256-bit identifier. A nominal type system plus mandatory boxing for recursive types decouples modules, so the pipeline runs parallel across cores with no query engine and no lock contention. Compilation walks flat arrays rather than pointer-chased trees.
The same source compiles to byte-identical MLIR whether it is built serially or in parallel. That is asserted in the test suite rather than hoped for — at benchmark scale, at one thread, at four, and with the thread pool taken off the path entirely, plus a corpus recompiled in fresh processes so each run gets its own hash seed. The claim is about the MLIR the frontend emits; everything downstream of it belongs to LLVM.
Backends
- CPU (x86-64, arm64): MLIR → LLVM IR → native, AOT or JIT
- NVIDIA GPU: MLIR → NVVM → PTX → SASS
- Apple AMX / ANE: CoreML primitive dispatch via plugin
- Distributed: Manifest-driven remote regions over a wire protocol
Vendors extend the compiler through MLIR pass plugins rather than by patching it.
Where Vx is the wrong tool
PyTorch users mutate architecture mid-loop, print a tensor shape, branch on it, and carry on. In Vx — ahead-of-time, data-oriented, statically regioned — that same dynamism takes real work.
Vx is the right language for the thing that must be correct and fast across ten kinds of silicon. It is not the right language for the thing you are still figuring out.
Toward provably private learning from federated data
Google's new federated learning system uses Trusted Execution Environments to provide verifiable, auditable privacy guarantees for decentralized data training.
Summary
Deep Dive
- Moves computation from individual devices to server-side TEE clusters.
- Uses Rekor (a transparency log) to publish access policies that auditors can inspect.
- Eliminates diurnal availability constraints by collecting device uploads before initiating training runs.
- Supports arbitrary Python-based training loops using Federated Language.
- Protects proprietary model architectures through runtime sideloading of serialized information.
- Gboard has deployed this system to achieve faster training times for next-word prediction.
Decoder
- Trusted Execution Environment (TEE): A secure area of a main processor that guarantees code and data loaded inside it are protected with respect to confidentiality and integrity.
- Differential Privacy (DP): A mathematical framework that adds noise to data so that the presence or absence of any single individual cannot be inferred from the output.
- Remote Attestation: A process where a system proves its software integrity to a remote challenger by providing cryptographic evidence.
Original Article
Full article content is not available for inline reading.
Multimodal Models Learn From Their Own Critiques
UniEvo-VL enables a multimodal model to act as its own teacher by using self-critiques as privileged information to improve performance during test-time.
Summary
Deep Dive
- Uses a self-distillation training recipe for multimodal generation tasks.
- Teacher-student roles are simulated within a single model by varying input context.
- Training minimizes per-state divergence between the student's and teacher's denoising diffusion distributions.
- Tested on benchmarks like GenEval and GenEval2 Soft-TIFA with significant improvements in image generation.
- Shows that models with stronger inherent 'judge' capabilities yield higher ceiling performance.
- Notes that self-improvement gains are not uniform across all tasks, such as text-rendering.
Decoder
- Self-distillation: A training technique where a model is trained using its own previous versions or its own predictions as the target.
- Privileged information: Information available to a teacher model during training that is not available to the student model at test time, used here to guide the learning process.
Original Article
UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement
Abstract:Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introduce UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback during test-time compute. Instead of relying on a separate, often larger, teacher, we leverage their self-critiques as privileged information and ask a single multimodal model to act as both teacher and student with different contexts. The student only sees the vanilla question, while the teacher conditions on the privileged critique. Then training minimizes the per-state divergence between their denoising diffusion distributions over the student's own sampling trajectories. Experiments demonstrate that UniEvo-VL improves the image generation capabilities of multimodal models, while maintaining their sensitivity to additional reflection information. Specifically, we build on top of the open-source Qwen-image-2512 and observe a significant performance gain from 0.747 to 0.808 on GenEval and from 32.97 to 35.53 on GenEval2 Soft-TIFA. Moreover, attempts with more powerful external critics (e.g., GPT5.6-Luna) show that multimodal models with strong judge capabilities can anticipate a higher self-evolving ceiling. Last but not least, mixed text-rendering outcomes show that our self-improvements may not be uniform across different tasks. Our study aims to shed light on the current hot recursive self-improvement research line to enhance the user experience when using multimodal models without external supervision or guidance.
Solving Open Research Problems Together
Meta researchers collaborated with mathematicians to solve six open research problems by using their Muse Spark AI models to draft proofs and generate counterexamples.
Summary
Deep Dive
- The research involved a 'human-in-the-loop' workflow where human mathematicians guided the AI, followed by a separate peer review process.
- Muse Spark successfully disproved several existing conjectures in group theory and non-associative algebra.
- The AI generated code in GAP, a computer algebra system, to find a counterexample in group theory.
- One study established a theoretical benchmark for the limits of fitting random Gaussian points to an ellipsoid in high dimensions.
- The team developed a bridge between number theory and p-adic string theory by using the model to connect distinct mathematical calculations.
- Meta emphasized that all papers explicitly credit the AI's contributions and denote which passages were drafted by the model.
Decoder
- Gaussian Ellipsoid Fitting: A statistical method used to find the best-fitting ellipsoid for a set of random data points in multi-dimensional space.
- Biharmonic Nonlinear Schrödinger Equation: A partial differential equation used in physics to model wave propagation, often studied in the context of laser and quantum phenomena.
- Semiabelian Group: A specific classification of mathematical groups, which are algebraic structures that represent symmetry.
- Monomial Property: A characteristic of groups where they can be represented by matrices of a certain simplified form.
- GAP (Groups, Algorithms, Programming): A software system for computational discrete algebra, particularly useful for research in group theory.
Original Article
Earlier this year, our models achieved gold-medal-level performance across five high-school Olympiad competitions in mathematics, physics, and chemistry. Those results inspired us to see whether AI could also help scientists solve open research problems.
Competition problems can be incredibly difficult, but those problems already have a solution. Open research is different. There is no answer key, no guarantee that an approach will work. Making progress means trying and retrying ideas, making new mistakes and resolving them, and sometimes going back to the beginning.
Over the past several months, we've partnered with mathematicians to explore problems across several areas of mathematics. They used both Muse Spark 1.1 and 1.2 in Thinking Mode through the regular meta.ai chat interface, with no custom research scaffold.
Open research takes time. Our goal here wasn't to mass-produce papers, but to empower researchers and help them develop mathematical insights that others can understand and build on. We’re proud to partner with the scientific community to help them advance research. These collaborations were done under the following principles:
- A team of mathematicians guided the research and worked with Muse Spark to explore ideas and develop the arguments.
- A second group of mathematicians then reviewed their work.
- Each paper clearly marks which passages were primarily drafted by researchers and which were drafted by AI.
- Each paper gives credit to the earlier research and mathematical ideas it builds on.
Today, we're sharing six such papers from that collaboration. Five present answers to previously open research questions.
After completing our work, we learned that other teams outside Meta had independently announced solutions to some of the same problems using different approaches. We recognize and appreciate their contributions and clearly acknowledge their work and how it relates to ours in the papers.
Probability: The Strict Threshold for Gaussian Ellipsoid Fitting
The paper answers a long-standing question about fitting random Gaussian points in high dimensions to an ellipsoid. To visualize the problem in low dimensions, imagine dots scattered on a plane and trying to draw an ellipse centered at a fixed point that passes through all of them. Our main result identifies a sharp threshold for the number of points that can be fitted in this way. Below the threshold, such an ellipsoid exists with high probability; above it, such an ellipsoid almost certainly does not exist. It gives researchers a theoretical benchmark for the limits of exact data fitting, while the behavior exactly at the threshold remains unresolved. Proof strategies were developed and revised with help from Muse Spark, under the guidance of Aykut Arslan, while four other mathematicians on our team checked and refined the arguments. We also acknowledge three independent concurrent works posted in August 2026. Misiakiewicz and Wen independently proved the Gaussian threshold. De la Cerda, Potechin, Tulsiani, and Xu established the Gaussian threshold up to a vanishing multiplicative factor. Koehler and Sohn obtained a broader universality result that includes the Gaussian threshold as a special case. These works and ours were developed independently and use different approaches.
Differential Equations: Finite-Time Blow-Up of Radial Negative-Energy Solutions for the Mass-Critical Biharmonic Nonlinear Schrödinger Equation
The paper answers a long-standing question about wave collapse in a model inspired by laser physics. To picture the problem, imagine a tug-of-war between one effect squeezing a wave inward and another spreading it out. Can the wave keep concentrating forever, or must it eventually collapse? For waves that are symmetric around a center and have negative energy, in two or more dimensions, the paper proves that collapse must happen within a finite time. This settles a question left open in 2015 and confirms a prediction from computer simulations in 2002 for this setting, giving researchers a clearer understanding of when wave collapse is unavoidable. Muse Spark helped work through calculations, test possible arguments, and revise the proof, while Dinh chose the problem and key proof ideas, and a separate pair of mathematicians reviewed and helped refine the work.
Group Theory: Semiabelian Groups Need Not Be Monomial
The paper disproves a conjecture about groups, mathematical structures used to describe symmetry. First proposed by M. Kida in 2024, it stated that every finite group with a property called "semiabelian" must also have another property called "monomial." Just as finding a black swan disproves the claim that all swans are white, one exception is enough to settle this question. The team found that exception in a group with 384 elements, showing that the two properties do not always go together. This helps mathematicians better understand how these groups are classified. Muse Spark generated the search program in GAP, a mathematical software system, that found the counterexample. Golich and her collaborators verified the result and completed the argument, with two other mathematicians reviewing the work. We also acknowledge the AI agent Nilradical, which reported a different counterexample to the same conjecture on September 16, 2026. Our result was developed independently.
Optimization: Tightness of the Cycle-Based Relaxation for Completed Length-Three Alpha-Cycles
The paper answers a question first posed by Del Pia and Khajavirad in 2026 about when a relaxation of a binary polynomial optimization problem captures the original exactly. To picture the structure, imagine three overlapping circles in a Venn diagram, each containing a set of yes-or-no decisions. For the family studied, the approximation is exact when each region shared by two circles (but not the third) contains exactly one decision. If any of those regions contains more than one, the approximation leaves a gap. This gives researchers a clear rule for when this particular simplification loses nothing and when it needs strengthening. Muse Spark helped reframe the problem using probabilities, identify a counterexample, and develop the proof strategy. Arslan and Trung Le checked the arguments, corrected gaps, and refined the final proof.
Arithmetic Physics: String Two-Point Function = Height Function on a Curve
Muse Spark helped researchers connect ideas from two fields (number theory and p-adic string theory), following a direction envisioned by Yuri Manin in the 1980s. Starting from a connection already known for the Tate curve, the model helped the team see how it could extend to a much broader class of curves. The paper shows that two calculations, developed in different mathematical languages, describe the same quantity. In simpler cases, the calculation comes down to counting how many initial digits the coordinates of two points share in a base-p number system. This gives researchers a concrete way to translate ideas and calculations between the two fields. Beyond helping identify the connection, Muse Spark generated candidate proofs and drafted three core technical sections, which the researchers then checked, corrected, and refined.
Non-Associative Algebra: On Solvable Evolution Algebras and a Conjecture by García-Martínez and Pérez-Rodríguez
The paper disproves a proposed conjecture for classifying evolution algebras, mathematical structures inspired by evolutionary biology. García-Martínez and Pérez-Rodríguez suggested a test for identifying a class called "solvable" algebras. The team found a small, three-dimensional example that passes their test but does not belong to that class. The paper goes further than finding a counterexample. It establishes an alternative rule that considers whole subspaces rather than individual elements, giving mathematicians a more accurate way to understand these structures. Working from the researchers' prompts, Muse Spark generated the counterexample and proposed alternative characterizations and proofs. Barei checked, refined, and rewrote the material. We also acknowledge independent work by Hu and Wen, who reported counterexamples to the same conjecture.
Building on this work
This work continues our investment in scientific research. We believe progress will require close collaboration between experts and AI, with results that are then carefully verified, transparent in how they were produced, and clearly communicated to the broader research community.
Acknowledgments
We’d like to thank the researchers who brought their questions and expertise to this collaboration, and everyone who checked proofs, and helped make the papers clearer. We're also grateful to the broader mathematics community whose work these results build on.
Pretraining on 50,000 Hours of Unlabeled Brain Data
Neuralink has trained its neural decoders using 50,000 hours of unlabeled brain data to improve device accuracy and longevity without extra calibration.
Summary
Deep Dive
- Neuralink decoders translate neural spikes into intent for device control.
- Calibration has historically been a mandatory, time-consuming step for every user.
- The team utilized 50,000 hours of raw, unlabeled data from ongoing clinical trials.
- Using unlabeled data allows the model to learn the fundamental statistical patterns of neural firing.
- Pretraining reduces the amount of labeled, calibrated data required to reach peak performance.
- This improvement directly increases the longevity and reliability of the decoder after initial setup.
Decoder
- Neural decoder: A machine learning model that translates raw electrical signals from the brain into digital control commands.
- Calibration: The process of mapping a specific user's brain patterns to desired device actions, usually requiring the user to imagine movements.
Original Article
Neuralink's decoders are machine learning models trained to convert neural activity into intended action. These decoders are unique to each user - all users have to perform calibration prior to using a Neuralink device. Neuralink has collected over 50,000 hours of unlabeled, freeform neural data since its clinical trials began. This post details how the company used this data to improve its models, making its decoders more accurate and long-lived.
Things that Apparently Cause Cancer
Harvard-led studies linking proximity to nuclear plants with cancer mortality are statistically flawed and can produce similar results for random landmarks like Costco.
Summary
Deep Dive
- Researchers at Harvard's T.H. Chan School of Public Health published multiple papers asserting an association between nuclear plant proximity and cancer.
- The studies used a 'proximity score' across radii of up to 200km, which Tilson and Stein argue is scientifically invalid for radiation exposure assessment.
- The authors replicated the methodology and applied it to other landmarks like Costco, private universities, and sports stadiums.
- The flawed model attributed millions of deaths to these locations, proving that the methodology itself is generating false correlations.
- The original researchers admitted the studies were ecological and did not prove causality, yet they were presented as a basis for public health policy.
- The failure to perform placebo testing or use a plausible causal pathway allows statistical noise to masquerade as evidence of harm.
Decoder
- Ecological study: A study that examines data at a population level rather than an individual level, often leading to potential 'ecological fallacies' where group-level associations are incorrectly attributed to individuals.
- Attributable fraction: A statistical metric used to estimate the proportion of cases that would be prevented if an exposure were removed.
Original Article
Things that Apparently Cause Cancer
What does this methodology measure anyway?
It would shock you to know that Costco, according to a new Harvard School of Public Health methodology, caused 120,687 cases of cancer mortality nationally every year, simply due to living in close proximity to their warehouses. Private colleges accounted for 83,782 more. Harvard itself is responsible for 5,842 cancer deaths. Living near a Major League Baseball field causes almost 19 times as many cancer deaths as living near a Major League Soccer pitch. These conclusions are obviously absurd, but they come from the exact same methodology that researchers at Harvard use to “prove” that nuclear power plants cause cancer in nearby communities.
Earlier this year, Adam Stein and I called out a series of Harvard studies showing that nuclear power plants were associated with higher rates of cancer incidence and cancer mortality as nonsense. Now that we’ve spent the last few months replicating the analysis and expanding it, we are able to show that no matter what landmark you use, their methodology will show an increased cancer risk and mortality. The studies simply do not prove anything about cancer incidence.
Let’s recap: In December 2025, researchers led by Yazan Alwadi at Harvard’s T.H. Chan School of Public Health published a paper in Environmental Health that claimed to find that cancer incidence increased for people living closer to nuclear power plants in Massachusetts. In March, the same researchers published an expanded nationwide study claiming a similar result—this time looking at cancer mortality rates, rather than incidence—in Nature Communications. This was followed by a paper in the Journal of Exposure Science & Environmental Epidemiology that looked at associations of lung, breast, and colon cancers. Most recently, a study of total mortality, not just cancer, was published in the European journal Environmental Epidemiology.
The papers construct a “proximity score” based on distance from nuclear plants, up to 120 km in Massachusetts and 200 km in the national studies (about 75 and 125 miles). Every ZIP code or county inside those radii is treated as exposed, with closer locations receiving higher weights. If a ZIP code or county is within range of multiple sites, then the effect is cumulative. That means that a county that is in close proximity to one nuclear power plant could have a smaller score than another county that is further away from any one power plant but within range of many.
From this proximity score, the authors run regressions testing the outcome (cancer incidence, cancer mortality, or total mortality) on the proximity of the county and a collection of covariates. From this, they use the statistical coefficients, construct the Relative Risk for each county, age, and sex group, and compute the Attributable Fraction, from which they derive the number of attributed deaths from the nuclear power plant from 2000 to 2018. The authors state that their methodology and results provide the scientific basis and public health justification for an expanded research program.
But proximity is not exposure. We have methods of measuring exposure for nuclear power plant workers. While the radiation exposure of nuclear workers will always be greater than or equal to that received by the surrounding public, most of the closely monitored US nuclear workforce receive no measurable annual dose. When workers are exposed to radiation, the average dose received is only 2 percent of the occupational limit. If operators and workers who are on-site at nuclear power plants receive an annual dose between zero and one-fiftieth of the occupational limit, how is it possible that residents 5, 10, 25, 50, 120, or 200 kilometers away would receive any measurable dose from the same plant?
In talks, the authors hedge that their papers are merely ecological studies that show association, but never prove causality. Ecological studies are used to understand the relationship between outcome and exposure at a population level. This leads us to ask: what would the mechanism of exposure be? Well, according to the authors, we can just ignore the broader literature and physics, and instead make up exposure pathways and mechanisms. The use of an “ecological study” allows a lot of leeway in terms of explaining the broader world.
Over the last several months, we have replicated the results of these papers. The authors supplied us with eight lines of code and answered a couple of questions about the covariates, which did not replicate the results. Most of our replication was done through first principles combined with trial and error. Once we were reasonably close to the results of the first national study on cancer mortality, we took the methodology and applied it to numerous other landmarks.
Ridiculous things you can “prove” caused cancer mortality
Everything causes cancer. Sounds cliché; maybe those California warning tags were right all along. But thanks to the methodology created by Alwadi et al., we can now prove that anything and everything causes cancer. Oh, sorry, that is too strong of an assertion. To use the authors’ words, since they have claimed their papers don’t prove causality, we can create an association between any physical landmark and cancer and, from that association, figure out how much cancer is attributable to that thing. Which is totally not the same as saying “that thing causes cancer.”
For instance, living near a private four-year university is associated with a 15-fold increase in cancer mortality when compared to living near a nuclear power plant.
Costco has the largest effect of all the locations we have tested. Over 2.2 million cancer deaths can be attributed to Costco; that’s more than 20% of all cancer deaths between 2000 and 2018. Hot dogs, bulk spices, and reasonably priced clothes come with a cost. But it goes to show that the authors’ choice of nuclear power plants wasn’t serious. Private 4-year colleges, including Harvard, are associated with the deaths of 1.59 million people, public 4-year colleges 626,212, and Superfund sites 895,258. Once you include the attributable deaths from state capitals at 132,389, we can account for 5.54 million cancer deaths, over one-half of all cancer deaths, between 2000 and 2018.
If we look at professional sports venues, the spread looks even more bizarre. Living near an NFL stadium was associated with 418,520 deaths, NBA arenas 471,794, MLB fields 944,769, NHL rinks 126,751, and MLS pitches 50,889. None of these sports have any kind of emissions that leave the playing field, except for the occasional home run or foul ball. The difference across the sports might be the kinds of concessions being sold. Baseball is right up there with Costco, selling copious amounts of hot dogs. Perhaps most surprising of all in the realm of professional sports is NASCAR, which accounted for 79,278 cancer deaths. One would expect the emissions, noise, smoking, and copious amounts of light beer to be more impactful than the emission-free ball and puck sports, but moving closer to a speedway might just save your life. With sports added to the mix, over 7.85 million cancer deaths can be attributed across a variety of landmarks; that’s 3/4s of all cancer deaths between 2000 and 2018.
The supposed 115,586 attributable deaths from nuclear power plants look like chump change to the real perpetrators. Harvard, when singled out, accounts for 111,003 deaths. That’s just one institution threatening every community within 200 km. Once again, we are left asking what the mechanism could be. Plus, our study only goes from 2000 to 2018; the school has been around since 1636, so the actual death toll could be much higher.
According to a webinar presentation given by Dr. Koutrakis, one of the coauthors, the supposed release of radioactive effluence from nuclear power plants comes from the occasional refueling of the reactors, when operators open the reactor core, and radioactive dust escapes from the fuel assemblies, past the containment vessel, through the outer shielding, and out into the world without ever being detected by the myriad number of dosimeters and radiological sensors throughout the facility. If we really are so terrible at measuring radionuclides (we are actually very good at it), then it is entirely plausible that Harvard could have unlawful access to special nuclear materials and the brains of Cambridge could have built a nuclear reactor under the squash courts without the proper licensure.
What does the methodology actually show?
Cancer is a scourge across the human race. You have a ~40% chance of getting cancer sometime in your life. Which is why we should take cancer studies very seriously. If one is showing a surprising result, we want to get to the bottom of it. All kinds of claims can be made with statistics, but if we want to make good policies based on evidence, it is worth spending time making sure the evidence upholds its claims.
In our original rebuttal, we remarked that the studies can’t prove their assertions because they lacked proper control. It seems they also neglected to do any placebo testing. They chose nuclear power plants because a story could be built around that framework. When the researchers got positive results across our nation’s nuclear power plants, they didn’t check what their shiny new methodology would do using other landmarks. This is their pitfall: by taking the easy way out—getting results and making up a story around those results without double-checking their method—the authors could have no idea that what they were actually capturing was the methodology itself. In our replication and expansion, we have shown that choosing a number of sites and landmarks from capitals to warehouse stores can all yield a positive result without any reasonable pathway for American citizens to become exposed to radiation, develop cancer, and then die. The variations are random noise, all sloping in a positive direction. Even when using randomized outcomes, the methodology outputs positive results.
The papers by Alwadi et al. don’t show a novel mechanism by which nuclear power plants meaningfully contribute to cancer mortality; instead, they show an association of data points across 400 km-wide circles. 48,519 square miles, about the size of Mississippi, is a massive area to claim any kind of exposure. Most studies measuring distance-based exposures look at much smaller distances, such as under 10 km (an area of 113 square miles). The methodology is blind, which can be a feature in research areas needing to avoid bias, but in this case it is a bug; it doesn’t understand radiation exposure or dose. All it understands are its inputs: coordinates, proximity, covariates, and cancer deaths. From these, it can give you a number and even a positive result, but it cannot explain why that result exists. It is still an ecological study, one with a sophisticated statistical technique, but not a very useful one.
Sophisticated statistics can produce extremely precise nonsense when the design fails to identify the causal mechanism. Bad science like this is especially dangerous when applied to something culturally plausible that gets a positive result—it is easy to twist a result into a compelling narrative, especially if it confirms one’s biases. Nuclear power plants create energy through radiation, and radiation causes cancer; therefore, nuclear power plants cause cancer.
It is especially telling that no matter what landmark we applied to the methodology, we have yet to get a negative result. There is, in fact, a real chance that you simply cannot get a negative result from this method.
The attributable number of deaths from this methodology is probably zero, but the attributable number of bad papers is at least four.
Why don't more developers “use the platform”?
Developers often avoid using native platform APIs because they prioritize framework familiarity, missing documentation, or the 'IKEA effect' of building their own tools.
Summary
Deep Dive
- The "use the platform" mantra advocates for native browser APIs over libraries for better performance and accessibility.
- Historical fragmentation (IE6-era) forced developers to rely on libraries like jQuery to mask inconsistency.
- Familiarity bias leads developers to npm packages even when CSS or native HTML elements have direct, superior solutions.
- Documentation disparity: platform docs were historically scattered or technical, whereas library READMEs were curated and accessible.
- The "IKEA effect": developers find it more fun and educational to build their own modal or storage logic rather than using simple native APIs.
- Over-engineering often happens because developers do not deeply understand the layers (e.g., CSS algorithms) underneath their code.
- AI impact: Models might bridge this gap by preferring native APIs, but they also risk churning out custom code that ignores platform-native helper functions.
Decoder
- IKEA effect: A psychological phenomenon where people place a disproportionately high value on products they partially created themselves.
- Polyfill: A piece of code used to provide modern functionality on older browsers that do not natively support the feature.
- Epicycles: A term used in software to describe increasingly complex workarounds added to a system to maintain an over-engineered design.
Original Article
For years, advocates for web standards, performance, and accessibility have implored web developers to “use the platform”. I’ve often been one of those advocates.
The argument is simple: why build something yourself, in JavaScript, when the browser can do it for you? Whatever you build, it’s likely to have poorer performance and worse usability than something the browser could just give you out-of-the-box.
I think it’s worth taking the other side, though, if for no other reason than to understand where the “platform-skeptic” developers are coming from. If “use the platform” is so obvious, then why do so many people seem to need convincing?
The most obvious reason is historical: for the longest time, browsers were playing catch-up with the ecosystem on top of them. Libraries like jQuery filled crucial gaps while browsers implemented equivalent APIs – and even then, you might have to wait for laggards like IE6 to age out before you could actually use them. Today, most browsers are evergreen (Safari is debatable, although ~7 times per year ain’t bad), but up until the 2020s or so, web developers had to deal with a decidedly lumpy web. In that environment, rolling your own is a sensible choice.
Another reason is familiarity: when you’re used to looking for React components on npm, that’s what you tend to reach for, regardless of the problem at hand. If you search for “sticky positioning” on npm, there’s no package that says “just use CSS position: sticky, you dolt.”
And often, even with a robust standard, libraries on npm would fill a useful gap between framework ergonomics and the platform underneath it. I always found it intriguing that many React developers preferred to stick to JSX and React idioms – raw DOM APIs felt “icky” – but were perfectly happy to use lower-level libraries where raw DOM manipulations are common. For example, a virtual list library might happily use raw DOM APIs for pure performance, while exposing higher-level primitives that a novice React developer could better grasp. In a sense, the ecosystem of React components led to a natural division of labor where those with more expertise packaged up unfamiliar platform APIs in a more familiar form factor.
Some of this effect was also driven by documentation. Many npm packages have lovingly detailed READMEs or websites with examples, tutorials, and screenshots. Whereas until MDN became cemented as the go-to place for web documentation (with web.dev as Google’s more future-facing arm), documentation for the web platform was scattered across blogs, StackOverflow, and sites like CSS Tricks. And many of these sites would just tell you to use a well-known library like jQuery or GreenSock!
If it were just about third-party libraries versus platform APIs, though, then I don’t think it could fully explain the antipathy toward “use the platform.” Developers who are lazy (or do I repeat myself?), and who just want a ready-made solution for whatever problem they’re facing, are unlikely to care whether that solution comes from npm, the browser, or copied off of someone’s random GitHub Gist. They want to solve their problem and move on. But there’s a different source of anti-“use the platform” that I want to explore.
For a certain type of developer, building things yourself is just more fun. And often the resulting code is easier to reason about, especially if you don’t have an encyclopedic knowledge of the web platform. And once you’ve built something, there can be a kind of IKEA effect where you want to maintain and tinker with your own homemade code.
As an example, let’s imagine you’re trying to build a modal dialog. You might visually understand how these are supposed to work: content appears on the screen, but the background is still visible although partially occluded, and maybe clicking outside the dialog dismisses it. So you might grab for position:absolute and z-index to position the dialog correctly – aha, but the background still scrolls, so you have to disable overflow on the body… And then if you understand something about accessibility, you realize you need to handle Esc to dismiss, and build a focus trap, and return focus to the element that launched the dialog, and…
For many developers, what I just described sounds like a nightmare (and a good way to build something that only half-works). But for many developers, this sounds like fun! Think of how much you learn as you start building this thing. And think about how you could start putting your own spin on it by adding animations, themes, an optional “close” button… Before you know it, you’ve built a library that’s ready to go on npm. That’s way more fun than just grabbing <dialog> and calling it a day – what a downer!
And for many of us, before APIs like <dialog> existed, this was how we learned the web platform! Many of the people who now advocate for “use the platform” were once themselves authors of polyfills, shims, and libraries. I know because I’m one myself! I spent years working on tooling for IndexedDB, WebSQL, and other browser storage APIs as part of my work on PouchDB, which eventually led to me feeling confident enough to sit in W3C standards meetings and even open issues and pull requests on the IndexedDB spec itself. Without the forcing function of a gap in the platform that needed to be filled, I don’t know if I would have found the interest or motivation to get to that level of expertise.
Of course, doing it yourself is not always an unalloyed good. Sometimes it just comes from pure ignorance. On the web platform in particular, I think one of the reasons there was such a proliferation of JavaScript solutions to problems that could be better solved by CSS, for example, is that many developers just didn’t take the time to deeply understand how CSS works.
And to be fair, CSS has historically been hard to understand! There’s a reason the site is called “CSS Tricks.” Things like the clear fix, floats, and the min-width: 0 trick are hardly intuitive. Rather than trying to understand CSS’s internal algorithm, it’s often much easier to just imagine the imperative logic you want and then express it in JavaScript. Plus, for years CSS didn’t have a straightforward way to express common patterns like line clamping, textarea resizing, scrollbar hiding, etc. So of course developers built it themselves using the tools they already understood.
I don’t even think this phenomenon of “avoiding the platform” is limited to the web. It can apply to any developer working on top of a platform they don’t fully understand. For example, at my work, we use ClickHouse for storing various kinds of analytics data. At one point, my coworker and I disagreed about how to store large JSON data in a column: he built a system for compressing it before storage, whereas I put the data in a separate key-value store and only inserted the key into ClickHouse. It turned out we were both wrong! ClickHouse automatically compresses data, and as a columnar data store you actually get better compression across rows if you just let ClickHouse handle it. And the separate key-value store was just a poor man’s version of what a columnar SELECT already does.
I only realized these things after actually taking the time to thoroughly read the ClickHouse docs and then write a benchmark to prove my hypothesis. In the end I was shocked that we had built something that was slower and clunkier than what the platform itself could give us out-of-the-box. The parallels with JavaScript and the web platform were hard to ignore.
I’m sure that if you’re a developer on iOS or Android, or someone building on top of a game engine, or really any kind of developer building on top of any platform layer, you probably have similar stories. There’s a reason that the stereotype of the grizzled senior engineer is someone who can take a junior’s baroque tangled mess of code and replace it with a single line. The more you learn, the more you’re able to wield your knowledge of how the entire system works end-to-end to create the smallest possible contribution to it (and thus reduce your maintenance burden in the long run).
I’ve been trying really hard not to talk about AI this entire post (because I’ve done it to death over the past year), but of course I can’t help but wonder how AI coding will impact this phenomenon. I actually have both an optimistic and a pessimistic take:
- Optimistic: because LLMs have an encyclopedic knowledge of whatever platform you’re working with, they can choose exactly the right platform API to deliver the experience the prompter asks for in vague English. And since this solution is likely faster and more correct than userland code, the agent will prefer it after rigorous testing and benchmarking. Furthermore, the “IKEA effect” goes away when developers are not actually writing the code themselves.
- Pessimistic: because LLMs seem to love duplicating code – for example, ignoring helper functions that already exist in favor of writing their own for the umpteenth time – the amount of custom, non-platform-idiomatic code will skyrocket. Developers won’t instruct their agents to test enough or to try enough alternatives, and will just commit the agent’s first draft. And because it’s always possible to add more epicycles, the agents will continue to iterate on over-engineered solutions that never should have existed in the first place.
In my own use of AI coding, I’ve seen both phenomena happen. I’d like to think that as models and coding harnesses get better we’ll start to veer more towards the optimistic outcome, but I can’t say for sure.
In any case, these are my longwinded and somewhat conflicting thoughts on “use the platform.” As a mantra I love it, because it succinctly captures a feeling I have when I’m looking at some overwrought pile of spaghetti code and thinking how much better and more elegant it would be if the author just understood the layers beneath them a bit better. At the same time, I’ve been that author, and I’ve felt the joy of building such beautiful, messy code (beautiful to me, anyway), so I think it’s worth understanding where such developers are coming from. For that reason, I’m sure we’ll be hearing “use the platform” for as long as there are platforms.
We're going to need default hard budget caps on pretty much everything
As autonomous agents begin to control billing APIs, hard budget caps must become a default requirement to prevent catastrophic financial runaway.
Summary
Original Article
We’re going to need default hard budget caps on pretty much everything
Here’s a product feature which the world is going to need a whole lot more of over the coming months and years: default hard budget caps. I’m talking about the feature of pay-by-usage services and APIs that lets you say “after $X/month, cut this thing off and return errors”. These need to be hard limits. Soft caps, “after $X/month, send me a warning email”, will not cut it.
Coding agents, and personal agents (coding agents wrapped in a less threatening UI), greatly reduce the friction of spinning up code that can do useful things. Sometimes those things cost money—calls to paid APIs, or hosted web applications, or systems that can bill for additional storage and compute.
Nobody wants to wake up to an email sent at midnight warning about a budget limit and find that, while they slept, their rogue service had consumed several hundred (or several thousand) more dollars of usage.
An argument against this is that businesses don’t want their hosted applications to start throwing errors because some budget was exceeded. I expect that most businesses and individuals would prefer errors to a surprise $10,000+ bill.
I think hard budget caps need to be the default. If someone wants to live dangerously they should be able to do that, but it needs to be on an opt-in basis. Have a nice, clear checkbox somewhere prominent:
Remove the budget cap. My application will not be shut down if I exceed the configured budget limit, and I will be responsible for subsequent charges.
The service I most want to see this from is AWS. I’ve heard plenty of stories from people who refuse to use AWS for personal projects out of (justified) fear that a runaway service might bankrupt them. I’ve also heard stories from people who didn’t anticipate this and ended up seriously burned.
... and it turns out AWS finally launched spending limits a few weeks ago! From their announcement New AWS experience helps builders get started and ship faster on 16th September:
When you’re ready to upgrade to a paid plan, you can set a monthly spend limit for your project based on your usage patterns so that you stay within your budget. If a project’s usage reaches its spend limit, your project is paused for that month.
See also Create a spend limit in AWS Settings, though that page warns that “We’re currently releasing our new experience to a limited number of customers.” Here’s hoping that hits general availability for existing accounts soon.
Google Cloud launched a similar feature in July, called Spend Caps, which lets you “set a monthly financial cap on specific services within a project”. Looks like this is becoming a trend!
In an ideal world, our agents could help with this. It would be great if agents started biasing towards recommending providers with hard budget caps, and warning new and inexperienced builders against deploying applications using uncapped services that might get them into trouble.
Last Week in Kubernetes Development: Week Ending September 27, 2026
Kubernetes v1.38 enters its enhancements freeze with 89 features, while new security patches address StatefulSet and Windows node vulnerabilities.
Summary
Decoder
- Confused deputy attack: A security exploit where a privileged entity is tricked by a less-privileged user into performing unauthorized actions.
- DRA (Dynamic Resource Allocation): A Kubernetes framework that allows for more flexible management and scheduling of specialized hardware devices.
- NTLM coercion: A technique used to force a server to authenticate with an attacker-controlled UNC share, often used for credential harvesting.
Original Article
Week Ending September 27, 2026
Developer News
Two security advisories were published on September 24: CVE-2026-2270 (rated Medium, 5.9) describes a confused deputy attack in the StatefulSet controller that lets a user with namespace-scoped write permissions on StatefulSet and ControllerRevision objects create a cross-namespace pod, and CVE-2026-76654 (rated Medium, 5.8) describes an NTLM coercion vulnerability on Windows nodes when a pod’s subPath is a symbolic link to an attacker-controlled UNC network share; both are fixed in the September patch releases (v1.34.12, v1.35.9, v1.36.5, v1.37.1) covered in the Release Schedule below.
SIG Cluster Lifecycle has a leadership change: Vince Prignano (@vincepri) is stepping down as chair, and existing Tech Lead Fabrizio Pandini (@fabriziopandini) will assume the chair role in addition to his Tech Lead responsibilities. The change is under a one-week lazy consensus period.
The CfP for Kubecon Europe 2027 is already open, and closes October 11th. This includes Maintainer Track sessions and Lightning Talks for SIGs.
Election Update
Voting closes on October 2. If you are an Kubernetes Org Member, and have not cast your ballot yet, please vote right away.
Release Schedule
Next Deadline: Enhancements Freeze, September 29 (AoE) / September 30 at 12:00 UTC
The Kubernetes v1.38 release cycle heads into its Enhancements Freeze this week. Following KEP Readiness, 11 enhancements were removed from the v1.38 milestone, leaving 89 tracked out of the 100 that had opted in. Any KEP that still wants to join v1.38 and isn’t already tracked now needs an approved exception. Please reach out in the #sig-release channel in Slack with any questions.
Kubernetes v1.38.0-alpha.1 has been built and pushed using Go 1.27.1. See the release notes and the GitHub release.
Patch releases v1.37.1, v1.36.5, v1.35.9, and v1.34.12 are now available. These releases bump to Go 1.26.8 and include a handful of bug fixes.
Featured PRs
142218: DRA ResourceSlice controller: optionally refuse to publish capacities and attributes with driver domain
pohly added an option for the DRA ResourceSlice controller to avoid publishing capacities and attributes associated with a driver domain. ResourceSlices are used by DRA drivers to publish device information that the scheduler uses during allocation and placement. This change gives drivers more control over which resource metadata is exposed through ResourceSlices while preserving the ability to publish the devices themselves. The work builds on the ResourceSlice and structured-parameter design described in KEP-4381.
142478: DRA Device Binding Conditions: graduate to GA
ttsuubasa graduated DRA Device Binding Conditions to General Availability under KEP-5007. Device Binding Conditions allow the scheduler to wait for network- or fabric-attached devices to become ready before binding a Pod. This avoids binding a workload to a node before the required device attachment has completed and allows failed attachment attempts to be reported back to scheduling. The graduation makes this workflow stable for DRA drivers and workloads that depend on devices requiring external preparation.
KEP of the Week
KEP-6361: Leader Election Recovery
Kubernetes controller managers use a Lease to ensure only one replica reconciles resources at a time. Today, if the leader cannot renew its Lease before RenewDeadline during a temporary API server or etcd outage, it exits. Restarting forces its controllers to rebuild their cached view of the cluster, extending the interruption after the API becomes available again.
The proposal adds a transport-level write gate to controller clients. When the manager stops leading, the gate rejects new writes and cancels writes in flight while reads and watches continue, keeping caches warm. The election loop keeps trying to renew the same Lease; if it succeeds before another replica takes over, the gate reopens and reconciliation resumes without a restart. A configurable recovery deadline can limit these attempts, and the former leader exits if another replica takes over. The Lease API and election protocol remain unchanged.
alvaroaleman, jpbetz, and michaelasp authored the KEP with SIG API Machinery. Proposed on September 14, it was merged on September 25. The enhancement issue tracks the work, and discussion is in the SIG API Machinery thread.
KEP 6361 is implementable and targets Alpha in Kubernetes v1.38 behind the disabled-by-default LeaderElectionRecovery feature gate.
Other Merges
- Fixed a regression in v1.38 where resource quantities written by a v1.37 or older apiserver with a decimal exponent outside the int32 range could not be decoded, failing reads of the whole collection.
- validation-gen:
+k8s:minimumand+k8s:maximumsupporttime.Durationfields with quoted Go duration strings, such as+k8s:minimum="1s". Integer payloads ontime.Durationfields are now rejected. SelfSignedCertKeyOptionsink8s.io/client-go/util/certaccepts aKeyGenerator, allowing self-signed certificates to be generated with a key algorithm other than RSA. The default remains a 2048-bit RSA key.- Added resource version to pod binding API
- Fixed a bug where
NoExecutedevice taints (DeviceTaintRule) did not evict pods using DRA-backed extended resources (pod.Status.ExtendedResourceClaimStatus) - resource.Quantity: calling String or encoding a quantity is now guaranteed to not mutate the instance
- client-go: restrict CA key encipherment usage to RSA keys. Key encipherment usage is no longer included for non RSA keys in self signed CA certificates generated from client-go
- Fixed an issue where a Job recreated with the same name as a recently deleted Job could remain unreconciled because it inherited pending Pod expectations from the previous Job
- kubectl describe node: per-pod resource percentages no longer print -9223372036854775808% on nodes that have not reported allocatable resources, and the “Allocated resources” totals now use the same resource accounting as the per-pod rows
- Added resource version to eviction api
- LimitRanger no longer re-validates the resource requests and limits a pod resize leaves unchanged. Only the values a resize changes are checked against the LimitRange, so an existing pod is no longer rejected on resize by a constraint it already violates
- Fixed false fractional-byte warnings for integer resource quantities larger than the int64 range
- Out-of-tree kube-scheduler plugins using the
NominatedPodsForNodemethod must now passlogger klog.Loggeras the first argument - Added support for using PKCS#10 signing requests signed with ML-DSA keys in CertificateSigningRequest, gated by the CertificateSigningRequestMLDSA feature-gate
- kubeadm: added support for the ML-DSA encryption algorithm. “ML-DSA-44”, “ML-DSA-65” and “ML-DSA-87” are now allowed values for ClusterConfiguration.EncryptionAlgorithm for new clusters using the v1beta4 and the still disabled (WIP) v1 API
- Fixed DRA kubelet plugin helper rolling updates for drivers with long valid names by shortening automatic DRA service socket paths when needed
- The documentation of
Node.status.volumesAttached[].devicePathnow states the platform-specific semantics: on Linux it is the host block-device node, on Windows it carries the CSI VolumeID - Fixed a race in client-go MutationCache indexed lookups that could temporarily hide a recently created replacement object
- Both
distribute-cpus-across-cores=trueandalign-by-socket=trueoptions were enabled, even when a single socket had sufficient capacity - Fixed a bug where kubelet logged
--manifest-url-headercredential values (e.g., Authorization tokens) to runtime logs at startup. Only header key names are now logged - kube-apiserver: added the alpha
ManagedFieldsOptOutfeature gate (off by default) - kube-apiserver: Requests to unknown
/apis/...paths and final delegation 404s now return a JSONStatusobject withContent-Type: application/jsoninstead of a plain-text “404 page not found” body - Fixed a bug where kubelet device manager may assign the same device ID to multiple pods simultaneously
Promotions
- DRADeviceTaints and DRADeviceTaintRules to GA
Deprecated
- Removed the deprecated
WindowsHostNetworkfeature gate (KEP-3503, withdrawn). The gate no longer guarded any code paths - Removed the
SeparateTaintEvictionControllerfeature gate fromkube-controller-manager. Remove this gate from existing--feature-gatesconfiguration before upgrading
Version Updates
- Bumped golang.org/x modules and mdlayher/socket to their current releases
- Bumped grpc to v1.84.0, lifted its pin, and bumped the etcd modules to v3.7.2
Subprojects and Dependency Updates
- containerd v2.4.1: Fixes CVE-2026-53493, leaked tasks after failed container start, and container creation when SELinux relabeling is unsupported; also v2.3.6, v2.2.9, v2.0.13, v1.7.36
- prometheus v3.15.0: Deprecates –log.level in favor of runtime.log_level, adds Unix Domain Socket scraping, stabilizes XOR2 float chunk encoding, and adds zstd scrape support
- etcd v3.7.2: new patch release with cherry-picks of bugfixes; also v3.6.15 and v3.5.34
- vertical-pod-autoscaler v1.8.0: new minor release brings improved unboosting for the alpha CPU Startup Boost feature, improved API validation, configurable status leases and several other changes
- vertical-pod-autoscaler v1.7.2: new patch release with various small bug fixes
Shoutouts
No shoutouts this week. Want to thank someone for special efforts to improve Kubernetes? Tag them in the #shoutouts channel.
Halving the Time: How Uber Eats Rebuilt Its Search Pipeline
Uber cut Eats search latency by 50% through massive architectural simplification and by utilizing AI coding agents to automate performance optimizations.
Summary
Deep Dive
- Reduced retrieval latency by removing low-value lexical paths and leveraging product-level embeddings.
- Optimized ranking by splitting hydration into parallel ranking and presentation phases.
- Implemented column-oriented data layouts for ad bids to minimize serialization overhead.
- Improved Go performance by switching pointer-heavy data models to stack-allocated value types.
- Used request hedging to mitigate tail latency in hydration dependencies.
- Utilized an agentic pipeline to identify bottlenecks in live profiles, draft PRs, and validate fixes with benchmarks.
Decoder
- Hydration: The process of populating a search result with metadata (pricing, images, stock status) after the initial ID-based retrieval.
- Request hedging: A strategy of sending identical requests to multiple backend instances and using the fastest successful response to reduce tail latency.
- Lexical retrieval: Searching based on text matches (e.g., keyword searches) as opposed to semantic or vector-based search.
Original Article
Full article content is not available for inline reading.
Impeccable (GitHub Repo)
Impeccable is a new AI-native design tool that provides 60 deterministic checks to catch common 'AI slop' patterns in frontend design.
Summary
Deep Dive
- Installs as a CLI tool or provider-native hook in agents like Cursor, Claude Code, and VS Code.
- Includes 60 deterministic (non-LLM) detector rules for UI quality control.
- Provides 24 commands (e.g., /critique, /polish, /harden) for iterating on frontend components.
- Manages 'product truth' via a centralized PRODUCT.md to prevent UI drift.
- Supports both comp-first (generative) and code-first workflows.
- CLI version allows for CI-ready scanning of static files or live URLs.
Decoder
- Design System: A collection of reusable components and guidelines that ensure consistency across a software product.
- Deterministic detector: A code-based check that uses hard logic instead of AI to find known issues (e.g., 'Does this color have sufficient contrast?').
Original Article
Full article content is not available for inline reading.
T3code (GitHub Repo)
T3 Code acts as a unified control surface for managing various AI developer agents like Claude Code, Cursor, and Grok across multiple operating systems.
Summary
Original Article
T3 Code
T3 Code is an "agent harness control surface". It enables control of the agents on your machine with a best-in-class mobile app (iOS, Android), web app and Electron-based desktop app.
Works with your subscriptions on Claude Code, Codex, Cursor, Grok Build, OpenCode, and Google Antigravity. If they're set up on your computer, T3 Code can control them.
"Wait, what are you selling me?"
Nothing. We built T3 Code because we wanted the best possible development experience with agents. We were inspired by existing solutions like the Codex desktop app, Conductor, Claude Desktop and Cursor Glass, but none met our bar.
We wanted something performant, remote-ready, and truly open. If we ever go the wrong direction, we want you to have everything you need to fork and build the editor that you want.
Installation
Warning
T3 Code currently supports Codex, Claude, Cursor, Grok Build, OpenCode, and Antigravity. Install and authenticate at least one provider before use:
- Codex: install Codex CLI and run
codex login - Claude: install Claude Code and run
claude auth login - Cursor: install Cursor CLI and run
agent login - Grok Build: install Grok Build CLI and run
grok login - OpenCode: install OpenCode and run
opencode auth login - Antigravity: enable it in Settings, then use Install Antigravity and Sign in with Google. No CLI is required.
Command line
curl -fsSL https://t3.codes/install.sh | sh
On Windows, in PowerShell:
irm https://t3.codes/install.ps1 | iex
Then run t3 to start the server and open the local web app. t3 service install keeps it running in the background, t3 update moves to a newer release, and t3 --help has the full reference.
To try it once without installing, run npx t3@latest instead.
Desktop app
Install the latest version of the desktop app from GitHub Releases, or from your favorite package registry:
Windows (winget)
winget install T3Tools.T3Code
macOS (Homebrew)
brew install --cask t3-code
Debian, Ubuntu (.deb)
Download the .deb from GitHub Releases, then:
sudo apt install ./T3-Code-*.deb
Arch Linux (AUR)
Stable:
yay -S t3code-bin
Nightly:
yay -S t3code-nightly-bin
Some notes
We are very very early in this project. Expect bugs.
We are (mostly) not accepting contributions yet. Small fixes may be considered. Big features will not be.
Documentation
Full docs live in docs/. There's no docs site yet.
If you REALLY want to contribute still.... read this first
Install vp
T3 Code uses Vite+ so you'll need to install the global vp command-line tool.
macOS / Linux
curl -fsSL https://vite.plus | bash
Windows
irm https://vite.plus/ps1 | iex
Checkout their getting started guide for more information: https://viteplus.dev/guide/
Install dependencies
vp i
Have a feature request? Start an Ideas discussion.
Need support? Join the Discord.
Rebalancer (GitHub Repo)
Meta open-sourced Rebalancer, a high-performance assignment solver library designed for hyperscale resource allocation and task scheduling.
Summary
Deep Dive
- Solver Architectures: Includes a C++ multi-threaded local search solver and support for external MIP solvers.
- Capability: Handles assignment problems involving capacity limits and balancing objectives.
- Debugging: Comes with a bundled Next.js web interface (Explorer) to visualize constraint violations and solver performance.
- Usage: Ideal for use cases like traffic routing, host migration, and workload distribution.
- Academic Basis: Backed by research presented at OSDI 2024.
Decoder
- Mixed-Integer Programming (MIP): A mathematical optimization technique used to find the best possible solution to a problem where some decision variables are restricted to integers.
- Local Search: An algorithm that starts with a feasible solution and iteratively makes small moves to improve the objective function.
Original Article
Rebalancer
Rebalancer is an assignment solver library that provides a generic and intuitive API for defining any assignment problem and the ability to optimize the assignment given a variety of implemented algorithms.
An assignment problem is any problem that can be defined as a decision of how to assign objects to containers, such that each object is assigned to exactly one container, given that it satisfies a set of constraints/rules and optimizes a set of objectives/goals.
The core solver is written in C++ and runs in a single process with multi-threaded parallelism. Currently, it can handle problems with ~1M objects and containers reasonably well. It's easily extensible to support new solving algorithms and expressions. Independent of the problem definition the user can choose from multiple solving algorithms. The most common are:
- Local search starts with an arbitrary assignment, and keeps performing simple moves (such as moving an object to a different container, swapping two objects, etc.) that are valid and improve the objective, until it can't find new improvements (or hits a user-defined moves limit or time limit). This solver is not guaranteed to find a global optimal solution, but it scales very well and can handle big problems.
- Optimal solver (mixed-integer programming) represents the problem as a set of mixed-integer programming expressions, and solves it using a generic external library (Rebalancer currently supports two commercial solvers, FICO Xpress and Gurobi as well as the open source solver HiGHS). These solvers will find optimal solutions given enough time, but they don't scale to handle huge problems well.
There is a finite (but easily extensible) set of predefined expressions that can be used to represent goals and constraints. A few examples of popular ones:
- Balance: make a given dimension balanced across containers. For example, say the objects are shards and containers are hosts, each shard has a given CPU utilization, and it is desired to distribute shards across hosts in a way that overall CPU utilization of all hosts is as similar as possible.
- Capacity: limit a dimension within containers. For example, say each shard (object) has a memory requirement (dimension), each host (container) has a memory capacity (dimension), and it is required that the sum of memory required by all shards in a host doesn't exceed the memory capacity of the host.
Users interact with Rebalancer via an interface which is available in C++ and Python.
At Meta, Rebalancer has been used for dozens of large-scale resource-allocation problems, including hardware and server allocation, ML training and inference placement, traffic routing, and load-balancing migrations. Its design, algorithms, and production experience are described in "Optimizing Resource Allocation in Hyperscale Datacenters: Scalability, Usability, and Experiences", published at OSDI 2024.
Quick Example
Four tasks, two hosts, one capacity constraint — host0 starts overloaded with three tasks and host1 has one. Rebalancer finds a balanced 2-2 assignment using local search or, optionally, a MIP solver backed by HiGHS, Gurobi, or FICO Xpress:
Python
from rebalancer import ProblemSolver
from rebalancer.specs import (
CapacitySpec, ConstraintSpec, LocalSearchSolverSpec,
MoveTypeSpec, SingleMoveTypeSpec, SwapMoveTypeSpec, SolverSpec,
)
solver = ProblemSolver(service_name="rebalancer", service_scope="example")
(solver
.set_object_name("task")
.set_container_name("host")
.set_assignment({"host0": ["task0", "task1", "task2"], "host1": ["task3"]})
.add_object_dimension("memory", {"task0": 10, "task1": 10, "task2": 10, "task3": 10})
.add_container_dimension("memory", {}, default_value=20.0)
.add_constraint(ConstraintSpec(capacitySpec=CapacitySpec(
name="memory_capacity", scope="host", dimension="memory")))
.add_solver(SolverSpec(localSearchSolverSpec=LocalSearchSolverSpec(
moveTypeList=[MoveTypeSpec(singleMoveTypeSpec=SingleMoveTypeSpec()),
MoveTypeSpec(swapMoveTypeSpec=SwapMoveTypeSpec())])))
)
solution = solver.solve()
print(solution["assignment"])
# → e.g. {'task0': 'host1', 'task1': 'host0', 'task2': 'host0', 'task3': 'host1'}
C++
auto solver = ProblemSolverFactory::makeProblemSolver("rebalancer", "example");
solver->setObjectName("task");
solver->setContainerName("host");
solver->setAssignment({
{"host0", {"task0", "task1", "task2"}},
{"host1", {"task3"}},
});
solver->addObjectDimension("memory",
{{"task0", 10}, {"task1", 10}, {"task2", 10}, {"task3", 10}});
solver->addContainerDimension("memory", {}, /*defaultValue=*/ 20.0);
CapacitySpec cap;
cap.name() = "memory_capacity"; cap.scope() = "host"; cap.dimension() = "memory";
solver->addConstraint(cap);
LocalSearchSolverSpec ls;
ls.moveTypeList() = {ProblemSolver::makeMoveTypeSpec(SingleMoveTypeSpec{}),
ProblemSolver::makeMoveTypeSpec(SwapMoveTypeSpec{})};
solver->addSolver(ls);
auto solution = solver->solve();
// solution.assignment() maps task → host
Installation
Build from Source
Ubuntu
# Prereqs
sudo apt install git pip python3-pex libfast-float-dev libgoogle-glog-dev clang-19 clang-tools-19 clang-format-19
# Build Thrift and Folly from source
git clone https://github.com/facebook/fbthrift.git
cd fbthrift/
./build/fbcode_builder/getdeps.py install-system-deps --recursive fbthrift
pip3 install pex --user
./build/fbcode_builder/getdeps.py --scratch-path ./installed --allow-system-packages build fbthrift
cd ..
# Clone
git clone https://github.com/facebook/rebalancer.git
# Configure and build
cd rebalancer/build
cmake -GNinja \
-DCMAKE_COLOR_DIAGNOSTICS=ON \
-DCMAKE_PREFIX_PATH="$HOME/fbthrift/installed/installed/folly/lib/cmake/folly;$HOME/fbthrift/installed/installed/fbthrift/lib/cmake/fbthrift;$HOME/fbthrift/installed/installed/fmt/lib/cmake/fmt" \
-DCMAKE_MODULE_PATH="$HOME/fbthrift/build/fbcode_builder/CMake" \
-DCMAKE_BUILD_TYPE=Debug ..
ninja
HiGHS (open source MIP solver)
Pick one of the following:
# Option 1: Install via conda
conda install conda-forge::highs
# Option 2: Install via pip
pip install highspy
# Option 3: Build from source
git clone https://github.com/ERGO-Code/HiGHS.git
cd HiGHS && mkdir build && cd build
cmake -GNinja .. && ninja
macOS
Prerequisite: Install Homebrew if you don't have it. After installing, open a new terminal so the
brewcommand is available.
# Install dependencies
brew install cmake ninja boost fmt folly googletest fbthrift
# Clone
git clone https://github.com/facebook/rebalancer.git
# Configure and build
cd rebalancer/build
cmake -GNinja \
-DCMAKE_COLOR_DIAGNOSTICS=ON \
-DCMAKE_PREFIX_PATH="/opt/homebrew/lib/cmake/folly;/opt/homebrew/lib/cmake/fbthrift;/opt/homebrew/lib/cmake/fmt" \
-DCMAKE_BUILD_TYPE=Debug ..
ninja
Fedora
sudo dnf install boost-devel.x86_64 fbthrift-devel.x86_64 glog-devel.x86_64 gtest-devel.x86_64 gmock-devel.x86_64 fmt-devel.x86_64
After Building
The default build produces the Rebalancer library. To build and run the bundled examples, pass -DTESTS=ON to CMake and rebuild:
# From rebalancer/build/
cmake -GNinja -DTESTS=ON -DCMAKE_BUILD_TYPE=Debug ..
ninja TasksOnHosts.exe
./TasksOnHosts.exe
This runs the tasks-on-hosts example — distributing tasks across hosts by memory capacity — and prints the resulting assignment to stdout.
More examples are in algopt/rebalancer/examples/ (shard allocation, web balancing, knapsack, and others). Each .cpp file in that tree is built as a standalone executable when -DTESTS=ON is set.
Install a Prebuilt Package
PyPI
pip install rebalancer
Debian / Ubuntu
# Primary (requires gh CLI)
gh release download --repo facebook/rebalancer --pattern "*.deb"
sudo dpkg -i rebalancer_*.deb
# Fallback (curl)
curl -sL $(curl -s https://api.github.com/repos/facebook/rebalancer/releases/latest \
| grep "browser_download_url.*amd64\.deb" | cut -d'"' -f4) -o rebalancer.deb
sudo dpkg -i rebalancer.deb
Fedora / RHEL
gh release download --repo facebook/rebalancer --pattern "*.rpm"
sudo rpm -i rebalancer-*.rpm
Rebalancer Explorer
Rebalancer Explorer is a web UI for inspecting and analyzing solver runs. It lets you browse a problem's objects, containers, constraints, and goals, and see how a solution scores against them.
Run with Docker Compose
The quickest way to try it is the bundled docker-compose.yml. From the repository root:
docker compose up --build
Then open http://localhost:3000.
Development Setup
Pre-commit hooks
This project uses pre-commit to run clang-format automatically before each commit.
pip install pre-commit
pre-commit install
Notes on Contributing
A complexity of contributing to rebalancer is that it must compile both on Meta's build infrastructure as well as in the open source world. This dual requirement has led to a somewhat strange CMake design where CMake searches the entire directory tree for files it can build and then classifies them as library files, tests, benchmarks, or other executables.
Citing Rebalancer
If you use Rebalancer in your research, please cite the following paper:
@inproceedings {298719,
author = {Neeraj Kumar and Pol Mauri Ruiz and Vijay Menon and Igor Kabiljo and Mayank Pundir and Andrew Newell and Daniel Lee and Liyuan Wang and Chunqiang Tang},
title = {Optimizing Resource Allocation in Hyperscale Datacenters: Scalability, Usability, and Experiences},
booktitle = {18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24)},
year = {2024},
isbn = {978-1-939133-40-3},
address = {Santa Clara, CA},
pages = {507--528},
url = {https://www.usenix.org/conference/osdi24/presentation/kumar},
publisher = {USENIX Association},
month = jul
}
License
Rebalancer is licensed under the Apache 2.0 License.
Introducing Web Search API via AI Gateway
Cloudflare integrated web search capabilities into its AI Gateway, allowing developers to ground agent responses in real-time internet data.
Summary
Decoder
- AI Gateway: A middleware proxy for AI applications that adds caching, observability, rate-limiting, and security between an application and LLM providers.
- Grounding: Providing an AI model with external, verified data to reduce hallucinations and improve accuracy.
Original Article
Full article content is not available for inline reading.
Beyond Observability: Evolving Production Operations in the Age of AI
Industry experts from Netflix, Genesys, and Groundcover argue that agentic operations require shifting from human-readable dashboards to high-cardinality machine-readable telemetry.
Summary
Deep Dive
- Accountability: Human operators remain accountable for outcomes, even when delegating decisions to agents.
- Context: Agents require organization-specific context (topology/metadata) to avoid making business-critical errors.
- Telemetry: Shift from visual graphs to high-cardinality, machine-interpretable logs.
- Verification: Emphasize canary-first deployment and sandbox testing for agentic actions.
- Emerging Reality: System behavior is increasingly non-deterministic; resilience is probabilistic and must be continuously tested via game days.
Decoder
- High-Cardinality: Data sets that contain many unique values (e.g., specific user IDs, request IDs), which are necessary for detailed troubleshooting.
- Canary Deployment: A release strategy where new software is rolled out to a small subset of infrastructure before a full rollout.
- MCP (Model Context Protocol): An open standard for connecting AI assistants to data sources and development tools.
Original Article
Full article content is not available for inline reading.
How we extended Apache DataFusion to execute one query across many machines
Datadog open-sourced Distributed DataFusion, allowing developers to execute a single Apache DataFusion query across multiple machines to handle heavy analytical workloads.
Summary
Deep Dive
- Distributed DataFusion extends physical query plans across nodes rather than just local CPUs.
- Uses Apache Arrow Flight for low-latency network data transfer between workers.
- Implements stages to manage network boundaries and shuffles for group-by aggregations.
- Includes a 'NetworkCoalesceExec' node to gather results from distributed workers into a final single stream.
- Designed to avoid distribution overhead for lightweight queries by opting for local execution when stats suggest it is faster.
- The system is written in Rust and integrates with existing Substrait-based planning pipelines.
Decoder
- Apache DataFusion: An extensible, high-performance query engine written in Rust that uses Apache Arrow as its in-memory data format.
- Substrait: A standardized intermediate representation for relational algebra that allows different query engines to share the same logical query plans.
- Predicate pushdown: An optimization technique that moves filtering logic closer to the data source to reduce the amount of data processed.
Original Article
Full article content is not available for inline reading.
How LinkedIn extended ClickHouse from distributed tracing to metric discovery and analytics
LinkedIn consolidated three legacy metric metadata systems into a single ClickHouse index, cutting memory usage by 80% while scaling to over 150,000 queries per minute.
Summary
Deep Dive
- The system uses a 'ReplacingMergeTree' table to handle deduplication of metrics data automatically during merges.
- Queries are routed through a gateway that performs dynamic SQL rewriting to maintain compatibility with existing RRDtool-based clients.
- Metadata is kept current via a continual change-capture pipeline from the time-series source of truth.
- Projections are used to optimize complex aggregations like service-to-unique-metric counts that legacy systems struggled to perform.
- The architecture utilizes a red-black deployment strategy across data centers to ensure query isolation and high availability.
Decoder
- OLAP (Online Analytical Processing): A category of software technology that enables the analysis of large volumes of data from multiple perspectives.
- RRD (Round Robin Database): A legacy system for storing time-series data using a fixed-size, cyclic buffer; often results in inflexible query patterns at scale.
- Cardinality: The number of unique elements in a set; high-cardinality data is difficult for many databases to index efficiently.
Original Article
Summary
- LinkedIn’s observability team runs distributed tracing and metric metadata discovery, and analytics on ClickHouse, after proving the database on tracing first.
- Even at 1% head sampling, that system generates around 0.8 trillion spans and 200 TB of uncompressed data per day.
- The metric metadata index consolidates three systems behind a single ClickHouse cluster, serving 150k+ queries/minute at 68ms average latency across 13B+ metrics.
- Migrating off a legacy stack cut memory use to roughly one-fifth and compute to about two-thirds, with headroom to double metric volume.
With over a billion members worldwide and thousands of interacting services, a single LinkedIn feed load or job search touches dozens of backend systems. When one of them slows down, finding the culprit means being able to see inside all of them.
That visibility falls to the company’s observability team, who own the full stack that keeps the platform measurable, from hosts, containers, services, and storage up through online, nearline, and offline workflows and real user monitoring, along with the ingestion, alerting, triaging, on-call, and visualization layers built on top. That stack covers every kind of observability signal, including metrics, logs, traces, events, profiles, and exceptions.
LinkedIn’s ClickHouse journey began two years ago with distributed tracing. The team put it into production across three data centers, building a system that handles roughly 800 billion spans a day and gives engineers a near-real-time troubleshooting surface.
Staff software engineer Jacob Zelek shared the next step in LinkedIn’s observability journey: how the team consolidated a legacy metric metadata system onto a single ClickHouse index, resulting in a system that costs less to run and answers questions the old stack couldn’t, with plenty of room to grow.
LinkedIn’s initial chapter with ClickHouse
In 2024, the team was ramping up an OpenTelemetry-based distributed tracing system that follows a request end-to-end, from a member’s app down through every backend service it touches. Even at 1% head sampling, that system generates around 0.8 trillion spans and 200 TB of uncompressed data per day. They needed storage that could keep up with that volume on the write side and stay fast on the read side.
Among their targets for the system, end-to-end ingestion had to complete in under 10 seconds, listing traces over a two-hour window had to return within 2.5 seconds at P99, and fetching a full trace by ID had to come back in 750 milliseconds or less. They also wanted a 7-day retention policy, tiered storage, customizable indexes, and query workload isolation. Finally, whatever solution they picked had to support KQL, one of LinkedIn’s preferred query languages.
Today, the deployment spans three data centers. Each has its own ClickHouse cluster of 22 shards fed by PubSub and a tier of ingesters. The team runs a red-black setup, writing to two clusters and reading from one, with the second ready to take over if the first has trouble, and a third handling validation. The schema is flattened, with span attributes carried alongside, and runs 3x local replication over a distributed table. Compression is around 5x. Each node takes in 350,000 spans a second at steady state, bursting past 1.2 million at max load.
The team reordered the table to put low-cardinality columns first, repartitioned on the field they actually query, and tightened Bloom filter false-positive rates on the columns where it was cheap to do so, bringing queries to run in under two seconds.
The next challenge: legacy metric metadata
Years ago, LinkedIn relied on a legacy time-series data logging and graphing system called RRDtool. However, the naming convention it introduced never left, and at LinkedIn’s scale and size, it had stayed.
Today, this metric is still referenced by an RRD, a single string that concatenates two to six standardized dimensions. Originally, users could target a single metric, but at some point the team introduced the ability to apply a regex across many of them to generate multiple series in one graph, or a separate graph per match.
That problem is compounded by the scale at which LinkedIn operates. The company has more than 13 billion metrics that are still referenced using RRD semantics, with roughly 30% daily turnover. Millions of graphs and alerts are defined this way, with dozens of tools and services still querying RRD metrics.
The challenge for LinkedIn’s observability team was letting people discover and analyze across all of it. For example, engineers might want to ask which hosts emit a given metric, which RRDs a service emits, which metrics match a pattern across every service, and how unique metric counts are trending for growth tracking.
The older stack had grown into three separate services behind a single gateway: a custom in-memory search index, a custom in-memory KV store, and Elasticsearch. By 2025, all three existing services had reached their scaling limits, managing three separate systems had become an operational burden, and even together they couldn’t serve some of the queries engineers were asking for.
Developing a ClickHouse ecosystem at LinkedIn
The team considered a number of options. Building yet another custom solution was unappealing, and sharded MySQL couldn’t scale reads. Elasticsearch was already in hand but didn't handle larger spanning aggregations very well.
ClickHouse, meanwhile, supported all existing queries. Benchmarked against the custom in-memory databases, most ran faster on ClickHouse. Its quotas and query logs let the team see which team or service was driving each query, and as a generic OLAP store holding every dimension, it allowed the team to support analytical queries they’d never thought of.
“The expertise we gained from using ClickHouse for tracing made onboarding a new project very easy. We’re starting to develop a ClickHouse ecosystem at LinkedIn.” — Jacob Zelek, Staff Software Engineer
The new ClickHouse-based architecture for legacy metric metadata
The team’s time-series database feeds ClickHouse through a continual change-capture system, while the gateway that fronted every metadata query now rewrites those queries into SQL and runs them against ClickHouse.
That gateway made the migration “transparent.” Every tool and service in the company already routed through it, so when the gateway stopped fanning out to three separate services and started writing SQL to ClickHouse instead, nothing downstream had to change.
Underneath, the table is a replicated ReplacingMergeTree, ordered to keep the discovery dimensions efficient, with deduplication keyed on a unique ID. Projections handle the heavier aggregations the old systems couldn’t, including service-to-metric counts, service-to-unique-RRD counts, and datacenter-to-host listings.
The results, and the road ahead
Today, the consolidated index serves more than 150,000 queries per minute across the fleet, with an average latency of 68 milliseconds. In terms of resources, the migration cut memory use to roughly one-fifth of what the previous three-system stack required, and reduced compute to about two-thirds. Because the metadata index is a single shard the team doesn’t have to think about resharding, it also has enough headroom to double the metric volume.
Now that every dimension is queryable, major customers are rewriting their old regex-heavy queries into more efficient native ones. This means the cluster may actually scale down over time, even as metric volume grows.
Context Is the New Code (50 minute video)
Patrick Debois argues that teams must manage AI coding context with the same rigor as application code, including testing, versioning, and CI/CD pipelines.
Summary
Deep Dive
- The Context Development Life Cycle (CDLC) parallels traditional DevOps with loops for generation, evaluation, distribution, and observation.
- Context should be linted and tested; just as code requires unit tests, context requires 'evals' to verify agent task completion.
- Registries for reusable agent skills are emerging, similar to npm or Maven, to share best practices across teams.
- Context drift is a significant risk; automated monitoring must detect when previously effective prompts fail as underlying models change.
- AI Bills of Materials (AI-BOM) are needed to track provenance of generated code and the context used to create it.
- Observability in production involves instrumenting agent actions to capture 'post-mortems' on agent decisions for future context refinement.
Decoder
- Context drift: The phenomenon where the quality of an AI agent's output decreases over time because the provided context or the underlying model behavior has changed.
- Agent hallucination: When an AI agent provides incorrect, nonsensical, or ungrounded responses confidently.
- MCP (Model Context Protocol): An open standard for connecting AI assistants to data sources and development tools in a unified way.
Original Article
Full article content is not available for inline reading.
LinkedIn Built Kafka. What Did It Change When Kafka Was No Longer Enough?
LinkedIn built 'Northguard' to replace Kafka's coupled partition model, allowing for more flexible storage and recovery at a scale of 32 trillion daily records.
Summary
Deep Dive
- Kafka partitions combine too many responsibilities, making cluster expansion and failure recovery bandwidth-intensive.
- Northguard uses a range-based keyspace that can be split dynamically, replacing static partition hashes.
- The system uses 'log striping' to distribute segments of a single log across disparate brokers, preventing the need to move large historical datasets during broker additions.
- A distributed metadata plane using Raft-backed 'vnodes' eliminates the bottleneck of a single controller.
- Xinfra provides a compatibility layer, enabling dual-writes during transitions so applications don't need to be updated simultaneously with backend migrations.
Decoder
- Raft: A consensus algorithm designed to be easy to understand; it ensures a cluster of nodes agrees on the same sequence of events.
- Broker: A server in a streaming system that stores and serves records for specific partitions.
- Gossip protocol: A peer-to-peer communication technique where nodes periodically exchange state information to maintain cluster membership.
Original Article
LinkedIn Built Kafka. What Did It Change When Kafka Was No Longer Enough?
Kafka started at LinkedIn around 2010, giving services a way to publish data, consume it independently and replay it when necessary. By 2025, LinkedIn reported processing more than 32 trillion records and 17 PB per day, across 400,000 topics, over 10,000 machines and roughly 150 clusters.
Operating that infrastructure required an expanding set of services for balancing, recovery and cluster management. LinkedIn's response was Northguard, a new log storage system, and Xinfra, a virtualization layer that lets applications use both Kafka and Northguard.
Northguard keeps producers, consumers, ordered logs, replication and replay. The changes happen underneath the surface: smaller replication units, a sharded control plane and topics that can move between storage systems.
A note on the evidence: This article examines LinkedIn's June 2025 design and deployment report. Northguard and Xinfra were internal, closed-source systems at the time; LinkedIn told InfoQ it would consider open sourcing them later. I could not find a public release for this article. The performance and reliability claims therefore come from LinkedIn, without a public implementation we can inspect or benchmark independently. Results on other workloads, hardware and Kafka configurations remain unverified.
Kafka partitions do a lot of jobs
A Kafka partition is an ordered log, but it also determines replication, broker placement and how consumers divide the work. Several responsibilities share the same boundary:
With twelve partitions, a topic has twelve ordered logs. Each has a replica assignment, a leader accepting writes and followers copying the data. In a conventional consumer group, partitions also determine the available processing parallelism.
LinkedIn's write-up identifies this coupling as a problem at its scale: a partition is a heavyweight replication unit, so moving or repairing one can affect a substantial amount of data.
Adding a broker means dealing with history
Suppose a cluster has three brokers, A, B and C, and is approaching its storage or throughput limit. We add Broker D. It starts empty; using its capacity for existing topics usually requires moving some partition replicas:
Those transfers compete with production traffic for network and disk bandwidth. LinkedIn built Cruise Control to automate partition reassignment, cluster balancing and recovery after failures. The transfer still has to happen, often under a throttle to protect the live workload.
Northguard reduces the amount of data bound to each replica assignment.
Northguard separates the log from its replication units
Northguard's hierarchy looks roughly like this:
A range represents a log covering a contiguous portion of the keyspace. It contains segments, each with its own sequence of records. An active segment accepts writes; a sealed segment is immutable. LinkedIn describes sealing a segment when it reaches 1 GB, has been active for over an hour, or loses a replica.
The segment is the unit of replication. Kafka also divides partition logs into segment files, but those files normally follow the partition's replica assignment. Northguard can give successive segments in the same range different replica sets:
The range preserves logical ordering while its segments are distributed across brokers. LinkedIn calls this log striping.
Don't move the old data
Now add Broker E. Northguard can include it in the replica sets of new segments:
Existing segments can stay where they are. As new segments are created, the allocator distributes new writes across the expanded cluster. This gives a new broker work without first copying historical logs onto it.
It also gives the allocator frequent opportunities to correct uneven placement. The effect develops as segments turn over; adding a broker does not instantly redistribute retained data. For me, this is the most useful change: capacity can start serving new traffic without a large reassignment job first.
Failure recovery changes as well
Suppose active Segment 42 has replicas on A, B and C, and Broker C fails. Northguard can seal that segment and create Segment 43 with a new replica set:
Writes continue on Segment 43 while sealed-segment replication restores the missing copy of Segment 42. Repairing the old data is separate from placing new writes.
LinkedIn credits this smaller replication unit with preserving produce availability without the consistency compromises it made in its Kafka configuration. That comparison concerns LinkedIn's deployment choices; it does not mean Kafka inherently requires sacrificing consistency whenever a broker fails.
Ranges replace indexed partitions
Kafka's partition count also affects key placement. With a mapping such as hash(key) % number_of_partitions, adding partitions can send subsequent records for a key to a different partition. That matters for ordering, stateful processing and keyed joins.
Northguard instead divides the keyspace into ranges that can split:
The range being split provides the synchronization point. Producers using unrelated ranges can continue. If R1 splits into R2 and R3, records in R1 precede records in either child. Merging ranges creates a corresponding ordering relationship between the parents and the merged range.
This also helps stream processing. Consider joining orders and customers by customer ID. If their key partitioning is aligned, matching records can be processed together. With ten partitions in one topic and sixteen in the other, a framework may need to reshuffle data first.
Northguard uses buddy-style splits and merges, giving ranges aligned keyspace boundaries across topics. LinkedIn chose this partly so processing frameworks could avoid such shuffles when joining streams on the same key.
The control plane had to change too
Kafka's move from ZooKeeper to KRaft changed metadata storage, but the cluster still has one logical metadata state machine and an active controller. LinkedIn describes that centralized control plane as a bottleneck at millions of partition replicas.
Northguard distributes metadata across vnodes. Each vnode is an independent Raft-backed state machine, led by a coordinator. Together, the vnodes cover a hash ring:
Topic metadata is assigned by topic name; range and segment metadata by range ID. LinkedIn calls this a Dynamically-Sharded Replicated State Machine, or DS-RSM, and describes deployments with 128 or more coordinators and state machines.
Each coordinator manages the metadata assigned to its vnode, including range splits and merges, segment state, replica assignments and topic deletion. These operations no longer all pass through one metadata leader.
Not every piece of state goes through Raft
Brokers also need connection addresses, broker attributes and enough information about the vnode ring to route requests. Northguard distributes this minimal global state through gossip, using SWIM for membership and failure detection.
Authoritative topic, range and segment metadata remains in the Raft-backed vnodes:
Gossip helps a broker find the coordinator responsible for a request. The coordinator's replicated state machine decides the metadata change. Membership dissemination and authoritative updates therefore have separate paths.
Northguard also revisits the disk path
Kafka relies on sequential I/O, the operating system page cache and zero-copy transfer where possible. Northguard's default segment store, fps store, uses a WAL, a file per segment, Direct I/O, a RocksDB sparse index and an application-managed cache.
The store batches appends, writes the WAL, appends records to segment files, calls fsync and updates the index. Direct I/O bypasses normal page-cache buffering; Northguard populates its own cache using knowledge of active consume streams.
This gives LinkedIn more explicit control over caching. It also keeps local disks in the write path. Segment storage is pluggable, but the published default uses small replication units on local disks to meet LinkedIn's latency requirements.
The durability comparison needs some care
LinkedIn reports the following configuration parameters for its two deployments:
| LinkedIn's deployment | Disk synchronization | Reported sync/batch thresholds |
|---|---|---|
| Kafka | Lazy sync | 10 seconds or 20,000 records |
| Northguard | fsync on all replicas before producer acknowledgement |
10 ms, 20,000 records or 10 MB |
LinkedIn says Northguard meets its Kafka performance SLOs with those stronger disk-persistence guarantees. The announcement does not provide a reproducible benchmark setup or latency percentiles for that comparison.
Kafka commonly relies on replication for acknowledged-write durability rather than requiring an fsync on every replica before each acknowledgement. The result reported here is that LinkedIn could afford the stronger sync requirement after redesigning its storage path. It does not establish an inherent durability limit in Kafka or a general performance advantage for Northguard.
Produce is a stream, not a sequence of independent RPCs
Metadata operations use ordinary request-response calls. Produce, consume and replication use sessionized streaming protocols.
A producer handshakes with the active segment leader, then pipelines appends within a window limiting data in flight. The broker can acknowledge multiple appends together, but only after their records are committed. Consume reverses the data flow, with the client controlling the window; replication uses the same streaming model.
Changing the protocol also creates a migration problem. Thousands of applications already use Kafka clients and know which physical cluster holds their topics. They need another abstraction before storage can change transparently.
Xinfra virtualizes Pub/Sub
Xinfra sits between applications and the underlying log storage:
Applications use Xinfra clients with a unified API. A topic can move between physical clusters or between Kafka and Northguard without changing its logical name. This depends on adopting the Xinfra client; existing Kafka clients do not gain that indirection automatically.
By the June 2025 announcement, LinkedIn reported that more than 90% of its applications were running Xinfra clients.
Topics have epochs
An Xinfra topic records changes to its physical storage through epochs:
The application can keep using payments while one epoch lives in Kafka and a later one lives in Northguard. Each epoch contains the shards representing that version of the topic. Clients obtain the physical mapping from Xinfra metadata instead of hardcoding it into application configuration.
Kafka to Northguard without stopping the world
Xinfra creates a new epoch in the destination. Producers migrate first and temporarily write to both systems:
Dual writes allow rollback if the migration fails. Consumers move later; once the transition completes, dual writes stop. Consumers can still read previous epochs until retention removes their data. LinkedIn states that applications continue operating and ordering guarantees are maintained throughout the transition.
By June 2025, LinkedIn reported migrating thousands of topics carrying trillions of records per day. Kafka and Northguard were still running side by side. These figures describe the rollout at publication time, rather than a complete replacement of Kafka.
Xinfra also hides the cluster boundary
Xinfra can group topics from different physical clusters into one virtual cluster:
A consumer can subscribe to those topics through the same logical environment even when they are stored in different systems. This lets LinkedIn support workloads larger than one physical cluster and change topic placement without making every application track the topology.
The complexity did not disappear
The application API becomes simpler, but Xinfra adds infrastructure that the platform team must operate:
- MySQL stores virtual and physical topic and cluster mappings.
- ZooKeeper handles membership, leadership and consumer-group allocation.
- Vitess stores consumer checkpoints, with Couchbase providing a low-latency cache.
For LinkedIn's roughly 150 clusters, centralizing this work can be worth the cost. With two Kafka clusters, maintaining the virtualization layer might cost more than the problem it solves.
Perhaps the partition was the real thing LinkedIn outgrew
Northguard and Xinfra separate responsibilities at three levels. Ranges define the logical keyspace and ordering; segments define replication and placement. Independent Raft groups distribute metadata management. Xinfra separates the topic applications use from the physical system storing it:
Reading this as LinkedIn outgrowing the partition is my interpretation. Its write-up identifies several concrete problems: coarse replication units, indexed partitions, balancing and centralized metadata. The common thread is how many operations depend on the same placement and ordering boundaries.
Kafka had already carried LinkedIn to an extraordinary scale. For a deployment with forty topics and six brokers, building Northguard and Xinfra would be a very expensive response to routine operational problems. The useful question is which costs in your own system come from the workload, and which come from coupling ordering, replication, placement and parallelism to one partition.
Introducing Clef: our open-source decision models, and new RL fine-tuning platform
Cloudflare is challenging specialized decision models by launching Clef and Clef-flash, open-source models optimized for programmatic, high-speed classification tasks.
Summary
Decoder
- Decision Model: A specialized machine learning model designed to classify inputs into bounded, structured outputs with probabilities, rather than generating open-ended text.
- System 1: A cognitive model term used in AI to describe fast, instinctive, and automatic processes as opposed to the slower, deliberate 'System 2' reasoning.
- Forward-deployed engineer (FDE): Engineers who work closely with customers or internal teams to build and integrate custom solutions on-site.
Original Article
Over the last few weeks, there has been lots of buzz around decision models such as Typesafe AI’s Jev System One model. While classifier models have been around for some time, Jev introduces a new decision model concept into the world of AI — a model that produces bounded structured outputs cheaply, quickly and consistently that can be added into a workflow when a decision is required. These models are capable enough to work over any set of inputs without constantly retraining the model to incorporate new classification categories. This contrasts with the world of Large Language Models (LLMs), which are largely non-deterministic, but are open-ended enough to reason and generate text and tool calls for agentic workloads.
Today, we’re releasing two Cloudflare-trained decision models, Clef and Clef-flash, hosted on Workers AI. Clef is currently the leader when evaluated against the Jev Decision Index, you can view full results on the live benchmark demo site. These models are smarter, faster, and fully Jev-API compatible, so you can experiment with these hosted models easily. We’re fully open-sourcing these models on Hugging Face under an Apache 2.0 license for you to run locally and experiment with yourselves.
Lastly, we’re excited to debut our new reinforcement learning (RL) product, which allows customers to fine-tune Clef to suit their use cases as well.
What is a decision model?
A decision model makes classifications to help agents decide how to act, based on certain probabilities. For example, you can pass in a customer support message (inputs) and ask if it is urgent and which team should handle it. A decision model will return typed answers with probabilities (outputs), which your code can use to route the ticket, trigger an escalation, or defer to a human. This means that a human does not necessarily need to be in the loop for agentic decisions anymore — agents can programmatically gather context, make decisions, and take actions on tasks, or defer to a human when needed.
Specifically at Cloudflare, we’ve been testing our new Clef model on our Threat Intelligence team to help us classify website domains. By giving a domain to Clef (with Browser Run) it can quickly identify categories that the domain falls under — for example, it might classify a domain with a 95% chance it is a fashion website, 85% ecommerce, <1% phishing, etc. This classification took our Clef model 2.2s to fetch, render, and classify the website. In contrast, our fastest general LLM gpt-oss-120b took 4.7s in the same workflow, and only returned two classifications. As a user, you can imagine how a 2x savings in latency and results can help us improve our threat intelligence workflows and be faster in identifying malicious or legitimate domains. Generalize this to any use case where you need to make quick programmatic decisions, and you unlock powerful agentic workflows that are able to autonomously decide, reason, and execute.
In music theory, a clef is a symbol placed at the beginning of a musical staff that assigns specific pitch names to the lines and spaces. A decision model is analogous to a music clef because it helps define the domain of the context and the subsequent notes (actions) that follow it. We chose Clef as the name of our family of decision models, as it serves similar purposes, and the CF hearkens to Cloudflare.
How is Clef different from other decision models?
Although the market is getting increasingly saturated with decision models, Clef has some unique properties that make us excited to release it to the public. First, it has a vision encoder so it’s able to take in images and classify visual content. This is different from Jev, which only does text classification today. Secondly, our model has a 64k context window (compared to Jev’s 32k), which allows users to squeeze more input state for the model to classify against.
Third, our model is accurate and powerful, scoring competitively against other decision models on the market across various quality benchmarks. We shortlisted some evaluations below that are important for decision-making as defined by the Jev Decision Index and scored some of the more popular models on the market for it.
| Benchmark | Clef | Clef-flash | Jev | DiffusionGemma Jev | Kev 9B | Laya |
|---|---|---|---|---|---|---|
| BFCL · case exact | 98.47 | 98.76 | 95.75 | 96.52 | 94.51 | 38.13 |
| ToolRet · nDCG@10 | 69.19 | 66.43 | 65.28 | 61.21 | 64.26 | 12.69 |
| API-Bank · accuracy | 91.93 | 93.11 | 88.19 | 83.66 | 56.30 | 11.41 |
| Home appliances · case exact | 82.95 | 97.73 | 52.27 | 42.05 | 25.00 | 0.00 |
| When2Call · accuracy | 72.37 | 65.58 | 80.97 | 75.44 | 49.62 | 11.94 |
| BANKING77 · macro-F1 | 94.20 | 90.93 | 79.74 | 74.28 | 84.83 | 14.29 |
| CLINC150+OOS · macro-F1 | 97.43 | 66.77 | 89.27 | 83.49 | 79.03 | 3.19 |
| BRIGHT · nDCG@10 | 45.91 | 39.26 | 47.52 | 42.94 | 38.53 | 19.90 |
| Amazon ESCI · macro-F1 | 57.48 | 57.39 | 55.21 | 53.37 | 49.22 | 24.40 |
| PhishNChips · accuracy | 79.60 | 75.05 | 62.55 | 85.35 | 50.75 | 50.15 |
We also ran benchmarks across Typesafe’s own eval suite and our Clef models fared well, beating Jev in 3 out of 4 areas. Notably, our Clef-flash performs exceptionally well, given how much faster it is.
| Workflow | Clef | Clef-flash | Jev |
|---|---|---|---|
| Invoice processing | 64.7 | 57.1 | 61.8 |
| Customer service | 76.3 | 77 | 76.0 |
| Security incidents | 62.9 | 61.7 | 61.7 |
| Agent trace observability | 68.5 | 69.8 | 71.6 |
On top of the latency benefits from the model itself, our Clef models are hosted on Workers AI. Because they are hosted on Cloudflare’s infrastructure, we’re able to take advantage of our GPUs at the edge, leading to low network latency and faster decisions.
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef \
-X POST \
-H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
-d '{
"model": "clef",
"state": "Checkout has been failing for every customer for the last hour.",
"questions": {
"urgent": { "type": "noul", "instructions": "Is this support request urgent?" },
"team": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"billing": "Payments, invoices, and refunds",
"technical": "Outages, errors, and configuration",
"sales": "Plans and upgrades"
}
},
"severity": {
"type": "score",
"instructions": "How severe is the customer impact?",
"criteria": ["No impact", "Minor", "Major", "Critical"]
}
}
}'
Clef also produces strictly typed outputs similar to Jev and is fully API-compatible, so you can make the swap extremely easily. The larger Clef model is your more powerful precision model, while the Clef-Flash model is great for latency-critical decisions.
How we trained Clef
Clef builds upon this concept, but uses a different base model as the backbone. We currently use Qwen as the base model and post-trained it to suit decision model use cases. During inference, Clef uses Qwen for a prefill-only pass, then scores the valid schema choices in parallel. The decision step is non-autoregressive, so there’s no intermediate text to generate token by token, making Clef significantly faster than autoregressive LLMs. Rather than generating intermediate text to produce structured answers, Clef and Clef-flash derive schema choices directly from internal backbone representations.
By freezing Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash, we jointly optimized the routing head alongside rank-256 low-rank adapters. Our post-training utilizes label-smoothed cross-entropy for valid schema outputs paired with a Brier loss to refine probability calibration. We also developed Reinforcement Learning for Calibrated Decisions (RLCD) to serve as a secondary optimization target.
How fine-tuning can extend the capabilities of Clef
We heard a lot of internal use cases that required fine-tuning our Clef model to be built into our agentic workflows at Cloudflare. These use cases are incredibly specific and we have had many years of labelled decisions that we could use to train a specific classifier. Because Cloudflare has more than 15 years of network data across different domains, we can fine-tune a model to fit these specific use cases which is more accurate and faster than our generic Clef model.
Our new RL service
We are offering a service to help customers fine-tune Clef to suit their workloads with our hands-on FDE team. From that, we’ll learn from our hands-on experiences to build a self-serve platform that customers can use to capture data, fine-tune, and redeploy the model, all on Cloudflare.
To do this, we leverage the primitives that we already have built on our Cloudflare platform:
- Cloudflare AI Gateway
- Cloudflare Workers AI
- Cloudflare Containers
- Trainer (NEW)
Try it out today
We’re excited to launch our first Cloudflare-trained ML model from the Workers AI team today. If you have specific use cases and are already customers of these products, we’d love to chat with you and be design partners as we experiment in this space.
Apache Iceberg 1.12.0: What's New, Breaking Changes, and Upgrade Guide
Apache Iceberg 1.12.0 brings production-ready support for advanced features like variant types and deletion vectors, but requires careful testing due to breaking API removals.
Summary
Decoder
- Deletion Vector: A storage optimization that tracks deleted rows without requiring a complete rewrite of the underlying data files.
- Variant type: A data structure that allows for storing semi-structured, schema-less data (like JSON) within a strongly typed table column.
Original Article
Apache Iceberg 1.12.0 broadens production support for major v3 capabilities, including Variant, geospatial types, row lineage, deletion vectors, encrypted metadata, and REST catalog operations. However, it also removes platform support and deprecated APIs, so treat this as a substantial upgrade and test compatibility carefully before rollout.
Why Your AI Agent Pipeline Is Slow (And How to Fix It Without Changing Models)
Restructuring a sequential guardrail pipeline into a parallel, short-circuiting architecture reduced AI agent latency by 86% without modifying the models themselves.
Summary
Deep Dive
- Baseline sequential latency: 13,874 ms.
- Optimized parallel latency: 1,828 ms.
- Accuracy held steady at 94%.
- Removed artificial sequential dependencies between independent policy and safety checks.
- Prioritized fast, deterministic rejects to save compute on expensive downstream checks.
- Used Terraform for infrastructure-as-code deployment of the pipeline architecture.
Decoder
- Execution topology: The structure and order in which tasks and checks are executed within a software pipeline.
- Short-circuiting: A programming pattern where a process terminates early once a result is determined, avoiding unnecessary downstream computation.
Original Article
A case study in diagnosing and fixing a real latency regression in a production guardrail pipeline for agentic AI on AWS Bedrock, and what it reveals about where LLM cost and latency actually go wrong.
Agentic AI systems, where a model plans, calls tools and acts autonomously against real infrastructure, need validation before every consequential action. Skipping that validation can allow an agent to take a destructive action because, for example, a prompt injection instructed it to.
But validation has a cost: Every check added to the request path is a delay the end user or downstream system has to wait through. The naive approach of running policy checks, prompt-injection detection and destructive-operation screening one after another is exactly what can turn a safe architecture into an unusably slow one.
This is the story of diagnosing that problem in a real five-layer guardrail architecture, what the fix looked like and the general lesson it holds for anyone building latency- or cost-sensitive LLM pipelines.
The Setup: Five Layers of Validation Before an Agent Acts
The architecture in question validates every action an AI agent proposes before execution against five layers: policy compliance, prompt-injection detection, destructive-operation screening, contextual anomaly detection and a final authorization gate.
It is built on AWS Bedrock and published as an open-source Terraform module so other teams can deploy the same pattern.
The first working version ran each layer sequentially: The agent’s proposed action passed through layer one, waited for a verdict, then moved to layer two and so on.
This is the obvious way to build it, and it’s also the version most teams ship first, because correctness comes before performance and sequential logic is the easiest to reason about and debug.
Initial baseline latency per validated action: 13,874 ms
That number matters because it’s not an edge case; it’s the baseline cost of doing validation the straightforward way. For any agent taking multiple actions in a session, that’s not a minor tax.
It’s the difference between an agent that feels responsive and one that visibly stalls on every step. It’s also exactly the kind of latency that can lead teams to disable a safety layer in production because it makes things too slow, which is its own failure mode.
The Diagnosis: Sequential Dependencies That Were Not Actually Dependencies
The instinct when a multi-stage pipeline is slow is to look for a single slow stage and optimize it. That wasn’t the actual problem here.
Profiling the pipeline showed that most of the five layers had no real dependency on each other’s output. Policy compliance and prompt-injection detection, for instance, don’t need to wait for each other to run; they are evaluating different, independent properties of the same proposed action.
They were only running sequentially because that’s how the code was structured, not because the logic required it.
The second finding was that the layers were not ordered by cost. Expensive checks were sometimes running before cheap ones that could have short-circuited the whole pipeline early.
If a cheap, fast check can already reject an action, there is no reason to run four more expensive checks afterward just to confirm what has already been decided.
Both of these are common patterns, not specific to this architecture. Pipelines accrete sequential structure by default because that’s the easiest thing to write, and the cost of restructuring only becomes obviously worth it once someone measures the baseline and sees a number like 13,874 ms.
The Fix: Parallelize What Is Independent, Short-Circuit What Is Cheap
The restructured pipeline made two changes:
Parallelizing independent checks: Layers with no dependency on each other’s output run concurrently instead of in sequence. The pipeline’s total latency for those layers becomes closer to the latency of the slowest concurrent check rather than the sum of all of them, making concurrency a major lever in a multi-stage validation pipeline.
Short-circuiting on cheap, high-confidence checks: Fast, deterministic checks, such as policy rule violations and known-bad-pattern matches, run first and can terminate the pipeline immediately on a clear reject, without paying the cost of more expensive downstream layers when the answer is already known.
Neither change touched what the pipeline actually validates or how strictly. The evaluation set, detection logic and accuracy bar stayed identical; this was a change to execution topology, not detection quality.
The Result
| Metric | Baseline (Sequential) | Restructured (Parallel + Short-Circuit) |
|---|---|---|
| Latency per action | 13,874 ms | 1,828 ms |
| Accuracy | 94% | 94% |
| False-negative rate, destructive operations | 0% | 0% |
That represents a 7.6x reduction in reported latency, with accuracy and the destructive-operation false-negative rate held constant across the evaluation set.
Additional details on the architecture and its open-source Terraform implementation are published for anyone who wants to verify or build on it.
The accuracy figures didn’t move because the fix was never about the detection logic; it was about how much of the pipeline’s wall-clock time was structural overhead versus actual work.
That distinction is easy to miss when a pipeline is slow and the reflex is to start tuning the model or the prompts, when the actual bottleneck may instead be in the way the surrounding pipeline is executed.
The General Lesson: Pipeline Topology Can Dominate LLM Latency
This case is specific, but the pattern is not. In systems that chain multiple LLM-adjacent steps, including validation layers, tool calls, retrieval-then-reasoning pipelines and multi-agent handoffs, the default execution shape is often sequential.
That happens because it is the natural way to write code, and sequential execution silently compounds every added step into more wall-clock latency.
The answer is not always simply to use a faster model. It is also worth asking:
- Audit for false dependencies: Steps that run one after another because of code structure, rather than because of a genuine data dependency, are the first thing to look for. If step B doesn’t actually need step A’s output, they may be able to run concurrently.
- Order by cost and confidence, not convenience: Cheap, high-confidence checks that can short-circuit a pipeline should run before expensive ones, not after.
- Measure before optimizing the model: It is tempting to reach for a smaller or faster model as the first lever. In pipelines with multiple stages, restructuring the topology through parallelization and short-circuiting can sometimes yield a larger, cheaper win without changing output quality.
The same discipline applies to cost, not just latency. A validation pipeline that executes five model-adjacent calls per action is paying for those calls’ token usage and associated API work regardless of whether they run sequentially or in parallel.
Parallelization primarily reduces wall-clock latency by overlapping independent work. Short-circuiting is what can reduce both latency and cost, because it avoids downstream calls entirely when an earlier check has already produced a decisive result.
That distinction forces an honest accounting of what actually needs to happen versus what is happening out of habit.
Where This Applies Beyond Guardrails
The specific numbers here come from a security and validation pipeline, but the pattern generalizes to any agentic system with multiple sequential LLM-adjacent steps: RAG pipelines that retrieve, then rerank, then generate; multi-agent systems where one agent’s output feeds another’s input unnecessarily; or evaluation harnesses that run checks one at a time when most of them are independent.
The question worth asking of any slow LLM pipeline isn’t only, “Which model should we swap in?” but also, “Which of these steps actually have to wait for each other, and which ones are just waiting out of habit?”
Further reading: Additional background on the guardrail architecture is available through its open-source Terraform implementation and in “Five Layers Between Your AI Agent and a Production Outage” on DZone.
Faster String Aggregations with Dimension Tables
DuckDB demonstrates that replacing high-cardinality string columns with integer surrogate keys can drastically speed up aggregations and lower memory pressure.
Summary
Deep Dive
- Strings waste CPU cycles on hashing and memory on duplication.
- Dimension tables map strings to narrow integer types (e.g., USMALLINT).
- Perfect hash aggregation becomes possible when keys fit in small, dense integer domains.
- Join the original labels back at the very end to avoid lookup overhead during the massive aggregation phase.
- For high-frequency updates, use an ANTI JOIN to identify and insert only new strings into the dimension table.
Decoder
- Cardinality: The number of unique values in a column.
- Perfect hash aggregate: A grouping optimization that uses an integer key directly as an array index, eliminating the need for complex hashing algorithms.
- Surrogate key: A unique, system-generated integer identifier used to represent a value from a different system.
Original Article
Faster String Aggregations with Dimension Tables
TL;DR: When a query groups on long, repeated strings, move the strings into a small dimension table with sorted, narrow integer keys. Aggregate on the keys, then join the strings back in at the very end. The query works as before, though on small fixed-width integers instead of variable-length text.
Analytical workloads are full of repeated strings: product names, country names, station names, user agents, category labels. Take DuckDB's public train services dataset, which has one row for every stop a Dutch railway train makes. The data comes from the open datasets published by the Rijden de Treinen (Are the trains running?) application. You can query it straight from its URL:
SELECT departure_time, station_name, type
FROM 'https://blobs.duckdb.org/train_services.parquet'
LIMIT 5;
The table has 380,959 rows but only 537 distinct station names, so each name is stored again in thousands of rows: Amsterdam Centraal (18 bytes) alone appears in 7,591 of them. Load it into a table to follow along, in the web shell or the DuckDB CLI:
CREATE TABLE train_services AS
FROM 'https://blobs.duckdb.org/train_services.parquet';
When you GROUP BY the station name, DuckDB has to process the full string for every row. For example, this query counts how many trains call at each station:
SELECT station_name, count(*) AS calls
FROM train_services
GROUP BY station_name;
The string work repeats for every row, even though there are only two distinct stations here, and only 537 across the whole table.
This post shows how to do that work on small integers instead, by giving each distinct string a number and looking up the strings only at the end.
Background
The approach is the star schema from data warehousing, used here for query performance. It came up when we looked at a user report of heavy memory use in a high-cardinality grouping.
That report turned out not to involve strings, but Richard Wesley pointed out that in his own work, building dimension tables and joining the strings back at the end made a large difference to string-heavy aggregations.
Why Strings Are Expensive to Group On
DuckDB computes a GROUP BY with a hash aggregation, keeping one hash table entry per group. The aggregation uses the hash of each group key to find the key's slot in a hash table, then compares the key with the one stored in that slot. For integers, this takes a few CPU instructions. For strings, the cost grows with the length of the string.
In DuckDB, a string value is a 16-byte structure. Strings of up to 12 bytes are stored inline. Longer strings store a 4-byte prefix plus a pointer to the actual characters. That design keeps short strings cheap, but many real-world labels are longer than 12 bytes.
For these longer strings, hashing reads every byte of every string in every row. A match can't be confirmed from the prefix alone, so DuckDB follows the pointer and compares the full string. When a new group appears, its string is copied into the hash table's own memory, which makes the table larger than it would be with fixed-width keys.
Copying strings also affects memory use. A wider hash table fits less well into CPU caches, and in larger-than-memory aggregations it reaches the memory limit sooner and has to spill more data to disk. DuckDB does apply dictionary encoding to string columns on disk, but that is a storage optimization: once a column is read into an aggregation, each group key is a full string again.
Integer keys avoid these costs, as their fixed width makes hashing and comparing them cheap. When the key range is small, DuckDB can avoid hashing altogether: if the statistics show that the keys fit in a small enough domain, the optimizer picks a perfect hash aggregate, controlled by the perfect_ht_threshold setting, which uses the key value directly as an index into an array.
Building a Dimension Table with Sorted, Narrow Keys
The examples build on the train_services table loaded above. Each row records one stop: a service_id, the date, the service type, the train_number, the station_code and station_name, and the departure_time and arrival_time. The repeated string we want to encode is station_name.
Step 1: Measure the Cardinality
First, find out how many distinct values the column has, since that count sets how narrow the key can be.
SELECT count(DISTINCT station_name) AS num_stations
FROM train_services;
This returns 537, and that count decides how wide the key needs to be. You want the narrowest integer type whose range still covers every distinct value, because a narrower key means fewer bytes per row in the fact table and a smaller entry in the hash table you group on.
The keys come from row_number(), which starts at 1 and only counts upward, so an unsigned type is the right fit, spending none of its range on negative values. UTINYINT (1 byte) holds up to 255 distinct values, USMALLINT (2 bytes) up to 65,535, and UINTEGER (4 bytes) up to about 4.3 billion.
Step 2: Build the Dimension Table
Next, assign each distinct string an integer key. The important detail is the ORDER BY station_name inside the window: the keys are assigned in string order.
CREATE OR REPLACE TABLE stations AS
SELECT
station_name,
(row_number() OVER (ORDER BY station_name))::USMALLINT AS station_id
FROM (SELECT DISTINCT station_name FROM train_services WHERE station_name IS NOT NULL);
Step 3: Store the Key in the Fact Table
Finally, replace the string column in the fact table with its key, a rewrite of the table that you pay for once. The dimension table leaves out NULL, so the LEFT JOIN keeps the rows without a station and gives them a NULL key.
CREATE OR REPLACE TABLE train_services_encoded AS
SELECT ts.* EXCLUDE (station_name), s.station_id
FROM train_services ts
LEFT JOIN stations s USING (station_name);
The fact table now stores the 2-byte USMALLINT key, so grouping on it is cheap. With a key range this small and dense, DuckDB can use a perfect hash aggregate, indexing directly by the key instead of hashing.
Querying the Encoded Table
With the key in place, the query aggregates on integers and only looks up the strings once the result is small.
WITH rollup AS (
SELECT
station_id,
date,
count(*) AS calls
FROM train_services_encoded
GROUP BY ALL
)
SELECT s.station_name, rollup.* EXCLUDE (station_id)
FROM rollup
LEFT JOIN stations s USING (station_id)
ORDER BY station_id, date;
For top-N queries, apply the LIMIT before the join, so that only ten strings are looked up:
WITH top_stations AS (
SELECT station_id, count(*) AS calls
FROM train_services_encoded
GROUP BY station_id
ORDER BY calls DESC
LIMIT 10
)
SELECT s.station_name, top_stations.calls
FROM top_stations
LEFT JOIN stations s USING (station_id)
ORDER BY calls DESC;
Variations
Comparison with ENUM
DuckDB's ENUM type is dictionary encoding built into the type system: values are stored as small integers, and DuckDB picks the integer width for you. The dimension table is the better choice when:
- New values keep arriving. An
ENUM's values are fixed when the type is created. A dimension table can grow. - You need attributes. A dimension table can carry extra columns, such as the station's city or the line it sits on.
- The data leaves DuckDB. Integer keys and a lookup table can be exported to Parquet or CSV and used in other tools.
Several String Columns
The pattern applies per column: build one dimension table for each high-repetition string column. Each key gets its own integer type, sized to the cardinality of its column, and a query only joins back the dimensions it needs.
Keeping the Dimension Up to Date
When new data arrives, add unseen strings with keys that continue after the current maximum using an ANTI JOIN to identify new records, then encode the new fact table rows and append them.
When the Pattern Does Not Help
Don't use it when:
- The strings are short. Values of up to 12 bytes are already stored inline.
- The column is nearly unique. If most values are distinct, the dimension table has almost as many rows as the fact table and little is saved.
- You query the data once. Building the dimension table and rewriting the fact table costs a full pass over the data.
- You don't aggregate on the column.
Conclusion
Grouping on repeated strings is expensive. By moving them into a small dimension table with sorted, narrow integer keys, DuckDB can aggregate on fixed-width integers and keep its hash tables compact. The strings come back in a final join against a result that is already small.
Jev plays Pokémon Red for $1.65 model bill (GitHub Repo)
An AI agent successfully played through Pokémon Red in under 38 hours for a total cost of $1.65 in model inference fees.
Summary
Deep Dive
- The system uses a Node-based Game Boy emulator to expose RAM state as JSON to the model.
- A harness manages basic pathfinding (A*) and menu navigation, while the LLM handles high-level decision making.
- The agent makes roughly 800 to 1,300 calls per hour, costing between $1.00 and $1.70 per 24-hour period.
- House rules were applied to ensure consistent speed and eliminate animation overhead.
- The codebase uses 'loop protection' to prevent the model from repeating failing choices indefinitely.
- It relies on a local disassembly of Pokémon Red to map memory addresses to game-readable symbols.
Decoder
- Harness: A software wrapper that provides an environment for an agent to perform tasks, handling the interface between the game's memory and the AI model.
Original Article
Jev Plays Pokémon Red
Jev, TypeSafe AI's decision model, plays Pokémon Red. There are no scripts or cheats: the harness reads the game's memory, lists the legal options with some facts about each, and Jev picks one.
Landing page: jev-pokemon.vercel.app
Result
Jev beat the game. The live stream ran on YouTube from Sep 25 to Sep 26, 2026 and has ended. The highlights are on the landing page.
| Total time | 37h 40m |
| Decisions made | 16,150 |
| Input tokens | ~39.2M |
| Total Jev cost | ~$1.65 |
| Median decision time | ~0.4s |
| Team wipes | 16 (14 at the Elite Four) |
| Elite Four attempts | 15 |
| Final team | Charizard 83, Graveler 62, Nidoqueen 45, Beedrill 44, Haunter 39, Primeape 29 |
You need your own legally obtained copy of Pokémon Red. No ROM is included or distributed here.
How it works
Game Boy emulator (Node) ──► read RAM ──► game state (map, party, battle, on-screen text)
▲ │
│ button presses ▼
harness mechanics ◄── Jev picks one ◄── legal options + facts
(A* pathfinding, menus) (type matchups, damage estimates,
"2 areas toward the objective", ...)
- Jev decides:
- every menu answer (names, starter, YES/NO, shop, heal, learn/forget moves)
- what to focus on (progress, heal, train, catch, shop, explore, team)
- which Pokémon to catch, and which to swap in and out of the team at the PC
- where to go and who to talk to
- every battle action (move, switch, item, ball, run)
- The harness never decides:
- It only reads memory and presses buttons. It never writes to game memory.
- Game knowledge is limited to the story milestones (what the next goal is and where it happens). Progress is checked against the game's real event flags.
- Hidden items aren't shown to Jev, since a human wouldn't know where they are.
- House rules:
- Text speed FAST and battle animations OFF, set once in the Options menu at boot.
- The player is named JEV and the rival BLUE.
- Every caught Pokémon gets a nickname: Jev spells it one letter at a time (A–Z or DONE). It has to be a made-up name, not a species name or a nickname already in use.
- Loop protection:
- Options that were already tried without anything changing get tagged, and Jev is told to try something new.
- If the same failing choice keeps coming back, the harness samples an alternative.
- As a last resort, it reloads the latest milestone checkpoint.
Setup
Requirements: macOS or Linux, Node 20+, and git.
git clone https://github.com/christianmat/jev-pokemon && cd jev-pokemon npm install npm run setup # installs rgbds (via brew), clones + builds pret/pokered, generates symbol data cp /path/to/your/pokered.gb roms/red.gb # your own dump of Pokémon Red (US/EU) cp .env.example .env # then fill it in
The ROM must be the US/EU release. Its SHA-1 is ea9bcae617fdf159b045185467ae58b2e4a48b9a, the same build pret/pokered produces. The harness reads RAM addresses from that disassembly, so other versions won't work.
.env
| Variable | What it does |
|---|---|
JEV_MODE |
gateway for real Jev via Vercel AI Gateway, mock for a free, dumb stand-in (the default) |
AI_GATEWAY_API_KEY |
Vercel AI Gateway key (VERCEL_AI_GATEWAY_API_KEY also works) |
JEV_MIN_INTERVAL_MS, JEV_MAX_PER_MIN |
Throttling (defaults 300 ms and 90 per minute) |
Run
npm start -- --speed 1 # real-time; local viewer at http://localhost:8787 npm start -- --speed 1 --resume # continue from the latest milestone checkpoint npm start -- --speed 1 --load <save-name> # load a specific save from saves/ npm run headless -- --steps 3000 # max speed, logs only npx tsx scripts/tools/save.ts # save the running game (writes saves/manual-*.json)
Logs
logs/jev-calls.jsonlhas every Jev call: the full state, the options with their facts, the probabilities, and the latency.logs/events.jsonlhas maps, milestones, saves and errors.
Cost
- Price: Jev costs $0.042 per million input tokens, and output is free. A typical call is about 1,200 tokens.
- Rate: at real-time speed the bot makes about 800–1,300 calls an hour. That comes to about $1–1.70 per 24 hours.
- Ceiling: the throttle's worst case, 90 calls a minute nonstop, is about $7 a day.
Project layout
| Path | What |
|---|---|
src/emu/ |
emulator wrapper (serverboy / GameBoy-Online core), save states, audio tap |
src/game/ |
ROM tables, RAM reader, collision grid + A*, region graph |
src/jev/ |
Jev client (throttle, cache, log), AI SDK gateway backend, mock |
src/agent/ |
mode detection, dialog and menus, overworld, battle, field moves and items |
src/knowledge/ |
story milestones |
src/server/ |
runner + local WebSocket viewer |
web/ |
local viewer page |
site/ |
public landing page (static; deploy with Vercel, root directory site) |
scripts/ |
setup, data generation, debug tools |
Deploying the landing page
site/ is plain static HTML with no build step:
- Create a Vercel project from this repo.
- Set Root Directory to
site. - Deploy.
Legal
- No ROM included. This repo doesn't contain or link to Pokémon Red, and you need your own legally obtained copy.
roms/is gitignored. - No Nintendo assets. Game data (symbols, names, maps, font) is read at runtime from your ROM. The symbol file is generated locally from the pret/pokered disassembly. That disassembly isn't redistributed here.
- Not affiliated. Pokémon is a trademark of Nintendo, Creatures Inc. and GAME FREAK inc. This is a fan experiment, not affiliated with or endorsed by Nintendo, Game Freak, The Pokémon Company or TypeSafe AI.
- License. GPL-2.0-or-later, because it builds on the GPL-licensed serverboy / GameBoy-Online emulator core.
Made by Christian Mathiesen at Frigade.
DuckDB 2.0 dynamic ATTACH statements
DuckDB 2.0 Alpha introduces support for dynamic ATTACH statements, allowing developers to inject environment variables directly into catalog connections.
Summary
Deep Dive
- Enables dynamic connection strings for external catalogs including Iceberg/Glue.
- Resolves a major workflow bottleneck by allowing secret management via standard environment variables.
- Simplifies the local-first lakehouse development pattern by reducing boilerplate code.
- Early reports from production pilots suggest high stability for S3 parquet processing and Iceberg table updates.
Decoder
- ATTACH: A DuckDB command used to connect external databases or data catalogs, enabling cross-database queries.
- Lakehouse: An architecture that combines the low-cost storage of data lakes with the management and schema enforcement of data warehouses.
Original Article
Test driving DuckDB 2.0 Alpha and so far I like it. A major enhancement that it brings is dynamic attach statements. This means that things such as usernames, passwords, and account numbers can now just be pulled from environment variables instead of having to hard code them in a SQL script or pivot to python. This script below pulls my AWS account number to attach an Iceberg Glue catalog. Previously, I had to hardcode the account number because attach statements only accepted literals. Some other things that I'm going to dig into soon are the enhancements to Quack and its official support for DuckLake. More to come... #duckdb
How has stability been? I don't know if I can wait another 3 weeks 😵
i run it in a production pilot. reading parquet in S3 from RDS-DMS and updating Iceberg tables with it. less than 1% fail rate for like 10 days in hourly runs.
Dynamic attach is the feature I have been waiting for. In practice the friction with DuckDB was never the engine, it was juggling catalogs and credentials every time you wanted to federate across two lakes. With Quack and official DuckLake support landing in the same release, the local-first lakehouse story finally feels complete. I would be curious how dynamic attach behaves with concurrent writers on the same Iceberg table, since that is where catalog consistency usually gets interesting.
Figma Motion Adds Custom Styles, Audio, Text Animations, and Lottie Export
Figma Motion moves to open beta, introducing audio integration, character-level text animation, and Lottie export functionality.
Summary
Decoder
- Lottie: A JSON-based animation file format that allows designers to ship high-quality animations across platforms without large file sizes or complex coding.
Original Article
Connect multiple GitHub organizations to Figma
You can now connect multiple GitHub organizations to a single Figma plan, making it easier to work with code across teams and projects. Keep your shared design systems in one place while connecting the GitHub organizations your teams use.
Available on Organization and Enterprise plans.
Stop blaming the model for slow AI. Use Doherty's threshold as a guideline.
Developers should rethink Doherty’s threshold for AI, splitting interactions into immediate acknowledgment and subsequent background processing to mask model latency.
Summary
Deep Dive
- Doherty Threshold: A rule stating that human-computer interaction should have a response time of less than 400ms to keep user productivity high.
- Interaction Splitting: Decouple the request acknowledgment (UI) from the execution (LLM inference).
- Skeleton Screens: UI elements that provide a visual placeholder during data loading.
- Streaming Responses: Showing text generation in real-time to bridge the gap between intent and completion.
- Optimistic UI: Updating the interface as if the request succeeded, then handling errors if they occur.
Decoder
- Doherty Threshold: The principle that if a computer responds within 400ms, the user remains focused and productive without feeling a delay.
Original Article
Generative AI changes how Doherty's 400-millisecond response threshold should be applied because a request now consists of two separate moments: acknowledging the user and producing the answer. Models may need several seconds to complete their work, but interfaces can immediately acknowledge input, show meaningful progress, and preserve the user's attention while generation continues. Designing this initial response as a strict performance requirement can make slow AI feel substantially faster without changing the model or inference speed.
Who is responsible for design tokens?
Design token governance works best when split by function: designers manage primitives, while engineering handles transforms and platform-specific outputs.
Summary
Deep Dive
- Primitive Tokens: The foundational values like color hex codes or spacing variables.
- Semantic Tokens: Tokens that describe purpose rather than value, e.g.,
button-bg-primaryinstead ofblue-500. - Governance Layer: The process of controlling how tokens are updated, versioned, and documented.
- Transformation Layer: Scripts that convert design JSON/YAML tokens into platform-specific code (e.g., Swift, CSS, Android XML).
- CI/CD Integration: Using pipelines to validate that changes in token definitions do not break downstream UI components.
Decoder
- Design Tokens: Named entities that store design decisions (colors, typography, spacing) to ensure consistency across different platforms and products.
Original Article
Design token ownership works better when responsibilities are divided according to the type of change rather than handed entirely to either designers or engineers. Designers can own primitive values, a mixed design-systems team can govern semantic tokens as a public API, and engineering can own transforms and platform-specific output. Automated checks can protect this structure by allowing value changes while flagging deleted or renamed semantic tokens that could break components and require migration.
Design Web Apps Visually, Sync with Code Instantly (Website)
Timeful offers a local-first visual editor for Vue and Tailwind that maintains two-way sync with your source code without proprietary file formats.
Summary
Deep Dive
- Local-first workflow: Projects live on the local file system as standard git repositories.
- Two-way sync: Changes in the visual canvas write directly to the code; manual code changes reflect instantly in the design view.
- Tech Stack: Specifically supports Vue 3 and Tailwind CSS v3.
- Git compatibility: Maintains clean diffs and preserves existing code formatting in script blocks.
- AI Compatibility: Since the project format is standard code, external AI agents can modify the codebase without needing specific plugin support for the design tool.
Decoder
- Two-way sync: A development process where changes made in a visual interface are automatically represented as code and vice versa without data loss or translation errors.
Original Article
Design and code, always in sync
Your design is saved as real production-ready code, with no export steps in between. Invisibly, in the background.
Build together
Designers and developers open the same Vue files. Designers edit the visual layout while developers build the logic. Every change gets saved to code.
Pixel-perfect in production
What you see on the canvas is exactly what gets published to production.
Local-first
Your project lives in a folder on your device. Not locked on a server you don't control.
No export steps, no AI credits
Your design is loaded from code and saved to code using predictable algorithms, with no steps in between - the code is your data format. You change the design, the code updates. Someone changes the code, your design updates.
Always up to date
Your project is a regular Git repository which you can share with the rest of your team. Pull the latest code changes from production and your design is updated instantly.
Share everything
Components, design tokens, even color space. All exactly the same for designers and developers: it's the same code.
Keep your design freedom
The same canvas, workflow, and tools you are used to.
Edit every style
Fine-grained control over every color, spacing, layout, and other styles.
Design tokens
Use Tailwind design tokens or select arbitrary styles for full freedom.
Freeform shapes
Drag shapes around and resize them freely on the canvas.
Auto-layout
Frames that resize and layout content automatically, with precise control over alignment and spacing.
Breakpoints and states
View breakpoints and states side-by-side and edit them separately.
Design together with AI
Your entire project is accessible locally on your device by any AI agent. All editable in standard Vue and Tailwind code.
Use any agent
Point any AI agent at the project folder on your device to make changes. Timeful automatically updates your design whenever changes are made to the code.
Your AI already knows Timeful
Timeful works with a regular Vue and Tailwind codebase. Everything is accessible as context, with no proprietery formats, frameworks, or servers in between.
Production-ready code, yours to publish anywhere
Timeful saves regular code to a folder on your device. Nothing is locked behind an export or server.
You own the code
The code for your designs is saved on your device as regular Vue components.
Publish anywhere
Publish or share your codebase anywhere. There is no custom code and no proprietary file formats that lock you in.
Clean diffs
Timeful makes the minimal possible edits. If you change a style, only 1 line of code changes. Even existing code formatting is kept.
Collaborate via Git
Push and pull your changes to Git to collaborate, or revert to a previous version.
Frequently Asked Questions
Which versions of Vue and Tailwind are supported?
Timeful supports Vue 3 using the composition API in Single-File-Components, as well as Tailwind CSS version 3.
Does Timeful work with build tools like Vite?
Yes, you can use any Vite setup and plugins in Timeful.
Does Timeful work with frameworks like Nuxt?
Yes, Timeful works with Nuxt.
What kind of Vue code does Timeful generate?
Timeful generates clean, formatted Vue code with Tailwind utility classes. It only modifies the template block and leaves your script block and logic untouched. Existing code formatting is kept and only the modified elements are reformatted.
Does Timeful work with React?
Not yet. The current version of Timeful works with Vue and Tailwind.
How does the two-way sync work with Git?
Timeful acts like a local visual IDE. When you make a change on the canvas, it writes directly to your local files. You commit and push those changes to your Git repository using your existing workflow.
Does Timeful store our source code on its servers?
No. Your code stays in your local repository and your version control system. Timeful runs locally, ensuring your intellectual property and source code remain completely secure.
Can I import my existing designs directly from Figma?
No. While Timeful features a Figma-like interface to reduce the learning curve, it operates directly on Vue files rather than importing static design layers. You can use a Figma plugin to export your design to Vue code, or rebuild components while using your Figma file as reference.
Does Timeful support custom Tailwind configurations and plugins?
Not yet. Timeful uses the default Tailwind configuration.
Anthropic to invest $100 million to train AI engineer talent
Anthropic is committing $100 million to the Claude Frontier Academy, aiming to train 10,000 enterprise AI engineers by late 2027.
Summary
Decoder
- Forward-deployed engineer (FDE): Software engineers who work directly on integrating complex technical solutions (like AI models) into the specific, often messy, operational workflows of enterprise clients.
Original Article
- Anthropic will invest $100 million in a program to train AI engineers.
- The program aims train 10,000 "frontier deployed engineers." by the end of 2027.
- As Anthropic pushes toward what's expected to be a historic IPO, the company is seeking to address a gap in available AI talent across industries.
Anthropic will invest $100 million into an academy for training AI talent, working with companies to integrate artificial intelligence into their operations.
The investment will fund a program called the Claude Frontier Academy, which aims to train 10,000 "frontier deployed engineers" drawn from companies that are part of Anthropic's Claude Partner Network by the end of 2027, the company said in a press release on Friday.
Engineers from Accenture, Morgan Stanley, Novo Nordisk and a number of top consulting firms are included in the first groups. The program was borne out of Anthropic's work partnering with businesses in implementing Claude, according to Anthropic.
"We hear our customers and partner organizations say, we need more people who can bring together familiarity and — with the enterprise tech and business context — and combine that with the highest level of AI fluency to solve the problems that need solving," Shambhavi Shambhavi, Anthropic's head of strategy and operations for partnerships, told CNBC.
Anthropic is in the midst of massive expansion as companies across the globe race to infuse its AI tools across their business units. The company is expected to hit the public market later this year, and could reportedly seek a $2 trillion market cap. Reuters reported earlier this week, citing a leaked copy of the IPO prospectus, that Anthropic generated almost $4.6 billion in revenue last year while racking up an operating loss of over $8 billion.
The training program will begin with an intensive "simulated enterprise deployment" with a graded assessment. Those who pass will enter a residency that follows a medical teaching model, with software engineers learning from instructors and practicing with casework before being assessed and credentialed. The first engineers are expected to be certified in early 2027.
"Claude Frontier Academy trains people the way our own engineers learn, and we want those who graduate to set the standard for how AI gets built inside a business," Steve Corfield, Anthropic's global head of business development and partnerships, said in the release.
The launch comes amid soaring demand for AI engineering expertise across multiple industries. Job postings for forward deployed engineers and other AI-related roles have surged in the finance world this year, according to data from Draup provided exclusively to CNBC.
Smaller models are the future of AI sovereignty
AI sovereignty should be defined by the practical agency to test, govern, and replace systems, rather than by achieving national technological autarky.
Summary
Decoder
- AI Autarky: The self-sufficient production of all AI technologies, including chips, training data, and foundation models, without reliance on foreign trade or technology.
Original Article
Australia does not need a national chatbot with an Australian flag stuck on the side. Nor do we become sovereign by signing a long-term contract for access to an American frontier model and calling that national capability.
If somebody else can change the price, alter the rules, withdraw the service or dictate its use, then the capability is rented. Rented capability is not sovereignty.
That does not mean Australia should attempt technological autarky. We cannot, and we should not. The real task is harder: retaining enough capability, visibility and choice that our institutions are not trapped when commercial, legal or geopolitical conditions change. That is what AI sovereignty should mean.
It is not a single model. It is not a procurement slogan. It is not a data centre announcement dressed up as strategy. It is the practical ability to choose, adapt, govern, replace and, when necessary, walk away.
Rented capability is not sovereignty
The first mistake in this debate is confusing access with control.
The leading US frontier models are often remarkable. They are useful for research, analysis, coding, writing, translation, customer support and a growing range of complex tasks. Australian organisations should use them where they make sense.
But these are services and they are controlled elsewhere.
When an organisation relies on a proprietary frontier model, it usually does not control the weights, the training choices, the policy settings, the processing environment, the update cycle or the commercial terms. It may not know exactly where data is handled. It may not be able to audit the system to the standard required for a high-accountability public decision. And if the provider changes the rules, the customer’s room to manoeuvre may be very small.
None of this makes frontier models bad. It makes them infrastructure with conditions attached. And infrastructure is political.
We have already learned this with cloud platforms, software ecosystems and social media. We are learning it again with AI, except this time the systems are likely to be embedded far more deeply into our administration, healthcare, education, research and critical infrastructure.
The more consequential the use, the more important the governance question becomes. Who gets to change the rules? Who decides what is permissible? Who can withdraw the service? Who can inspect the system? Who can turn it off?
If the answer is a company operating in another jurisdiction under another legal system and pursuing its own commercial interests, that is not a minor procurement detail. It is the issue.
The monopoly assumption is breaking
For much of the generative AI boom, the working assumption was that the United States would remain decisively ahead, and that the best systems would stay locked behind proprietary APIs. That assumption is starting to fray.
The 2026 Stanford AI Index found that the gap between the leading US and Chinese models had narrowed to low single digits. As of March 2026, Claude Opus 4.6 led Dola-Seed-2.0 Preview by 39 Arena points, or 2.7 percent.
The United States still holds enormous advantages in capital, compute, frontier labs and infrastructure. That matters, and it is not going away. But a small and fluctuating performance gap is a different strategic environment from one of overwhelming and durable US dominance.
Australia should not respond to that by swapping dependence on US systems for dependence on Chinese ones. That is not sovereignty. That is vendor diversification with a geopolitical flavour.
What matters is that the model landscape is becoming more plural. There are more capable systems, from more providers, under more varied licensing and deployment arrangements. That includes open-weight models that can sometimes be inspected, benchmarked, adapted and run in environments the user controls.
More models do not make Australia sovereign. They simply remove one excuse for remaining dependent.
Open weights are not a magic wand
There is now a lot of magical thinking attached to open-weight models. People hear that weights are available and leap to the conclusion that capability has somehow become local, accessible and sovereign by default. It has not.
Open-weight models matter because they widen the field of possible action. Stanford’s 2026 analysis found that the gap between the best open and closed models had narrowed substantially, and it identified Alibaba among the leading providers in a more tightly clustered frontier field.
That matters because capable institutions can do more than simply rent access through somebody else’s interface. They can inspect, test, adapt and in some cases operate models under their own governance settings.
But let’s not get carried away. Downloadable weights are not the same thing as usable capability.
A large model still needs compute (aka data centres), skilled people, secure infrastructure, evaluation, monitoring and money. It may still depend on foreign chips, foreign cloud capacity, foreign software frameworks and foreign research ecosystems.
Open weights are not sovereignty in a box. They are leverage.
And leverage is what gives institutions room to negotiate, contest and walk away. I have often argued that open institutions matter because they preserve the ability to resist concentration of power rather than simply accommodate it.
Small is often the strategic choice
The sovereignty conversation has been distorted by frontier-model theatre.
Most organisations do not need a model that can do everything. They need a system that can do a particular job, reliably, securely and at a cost they can sustain.
A council classifying planning documents does not need a trillion-parameter model. A health service answering staff questions from internal policy documents probably does not either. A regulator searching its own guidance certainly should not be sending sensitive material into an opaque external system by default.
For many of these jobs, a smaller model in a controlled environment is the better choice. Not because small is automatically safe, but because smaller systems are easier to host, test, bound, monitor and replace. That makes them easier to govern.
Smaller and specialised models can be:
- Hosted closer to sensitive data
- Adapted to Australian law, institutional language and defined workflows
- Evaluated more thoroughly against a narrower and better-understood set of risks
- Run at lower cost and with lower energy demand
- Operated with lower latency and better resilience
- Replaced without rebuilding the whole organisation around one vendor
The design principle should be simple: use the smallest model that can safely, reliably and economically perform the task, then justify every escalation in scale.
This is not an argument against frontier models. Some work will need them. But “use the largest model available” is not a strategy. It is a habit formed in an era when capability was scarce, concentrated and theatrically marketed.
Australia cannot go it alone
Let’s dispose of one fantasy right now. Australia is not going to achieve AI autarky.
We are not going to manufacture every advanced chip domestically, train every major foundation model from scratch, own every cloud layer or detach ourselves from global research and software ecosystems. Nor should we try.
Autarky would be ruinously expensive, technologically brittle and strategically foolish. It would replace one set of dependencies with a smaller, more fragile local monoculture. But rejecting autarky does not mean accepting helplessness.
The real alternative is managed interdependence.
That means understanding where power sits in the AI stack and deciding, consciously, which dependencies are tolerable and which are dangerous. It means knowing where the chips come from, where the model is hosted, who controls the updates, what the licence permits, which cloud services are indispensable and which skills are so scarce that they become strategic bottlenecks.
It also means having options before we need them.
An institution should not discover during an outage, contract dispute, export-control shock or geopolitical rupture that one external provider has quietly become the operating system for an essential public function.
That is not resilience. It is wishful thinking with an invoice attached.
Sovereignty is tested under pressure
Real AI sovereignty is not a claim about where a model was trained, where a company is headquartered or where a server happens to sit. It is a test of what happens under pressure.
Can you:
- Inspect the system?
- Independently test its performance, safety, privacy and security?
- Move the data?
- Shift workloads to another model?
- Keep operating if the provider changes the terms?
- Tell when the model has changed?
- Meet your legal, regulatory and public-accountability obligations?
- Turn it off?
If the answer to those questions is no, then you do not have sovereignty. You have a dependency that has not yet been tested.
This matters most in government, regulated sectors and critical infrastructure. We should be very cautious about building essential functions around systems that cannot be independently evaluated, meaningfully audited or credibly replaced.
A proprietary model may still be the right choice in some cases. There is nothing inherently wrong with buying a service. But the decision should be deliberate, proportionate to the risk, and backed by contractual, technical and operational safeguards.
The point is not to eliminate dependency. The point is to ensure that dependency does not become obedience.
No single national champion required
The choice is not between a small Australian model and a giant American or Chinese one. That is not how a serious AI strategy works. Australia needs a portfolio approach, we need:
- Routine, private or low-latency work: Small local or Australian-hosted model
- Regulated decision support: Domain-adapted models with defined evaluation, logging and human oversight
- Complex research, multimodal analysis or agentic work: Large open-weight model where infrastructure and governance justify it
- Exceptional general-purpose capability: Proprietary frontier model, used selectively with security, contractual and exit controls
- Nationally significant uses: Australian research, evaluation, operational expertise and secure infrastructure across the stack
This is not about picking a national champion. It is about avoiding a national dependency.
The more AI becomes part of essential systems, the more we need to be able to negotiate, contest and walk away from the companies that provide it. The most important policy task is not to produce another strategy document full of aspirational language about trustworthy or responsible AI.
It is to build the capacity to choose, and that means investing in:
- Australian-hosted inference and secure compute for sensitive workloads
- Public-sector capability to procure, test, monitor and govern AI systems
- Independent model evaluation for security, reliability, privacy, bias and cultural appropriateness
- Fine-tuning, deployment and operational expertise across government, universities and industry
- Open standards and interoperable architectures that reduce vendor lock-in
- Indigenous data governance and meaningful community authority over culturally sensitive data and AI use
- Practical requirements for portability, continuity and tested exit plans in high-consequence procurement
- The technical and legal ability to pause or shut down AI systems safely when required
This is not glamorous work. It will not produce a shiny national chatbot to unveil at a press conference. It will produce something much more useful: institutions that can use AI without surrendering their capacity to govern it.
The future of AI sovereignty is not one giant sovereign model. It is the ability to choose the right model for the job, understand the dependencies it creates, and change course when the provider, the contract or the geopolitical environment changes.
That is the capability Australia needs to build. Not independence from the world, but enough agency to remain a participant in it rather than merely a customer.
Muse, dots, Instinct, and the question every agent will be asked
World ID provides Proof of Human credentials for AI agents to prevent rate limiting, spam, and unauthorized autonomous actions.
Summary
Deep Dive
- Addresses friction where services block AI agents to prevent abuse or bot-driven traffic.
- Enables one-human-to-many-agents mapping while keeping agent activity uniquely attributable to one user.
- Zero-knowledge proofs preserve user privacy during the verification process.
- Already being tested by Okta, Vercel, Exa, and Browserbase for human-in-the-loop workflows.
- Provides a mechanism for websites to set agent query limits based on verified human ownership.
Decoder
- Zero-knowledge proof (ZKP): A cryptographic method that allows one party to prove to another that a statement is true without revealing the actual data.
Original Article
Muse, dots, Instinct, and the question every agent will be asked
TL;DR
AI agents such as Meta's Muse, OpenAI's dots, and Instinct put the power of personal assistants in the hands of consumers. As they gain widespread adoption, websites and services will ask which agent is making a request and whether a real person originated the action. World ID lets an agent answer the second question while the person’s personal information stays private.
What is Muse?
Personal AI agents that act on behalf of individuals are becoming mainstream. Last month, Meta launched Muse, an agent that performs tasks like sending emails, booking travel, completing forms, and making purchases.
What are dots?
Three weeks after Muse arrived, OpenAI launched dots: always-on AI agents that operate in the background on cloud-based computers. They can connect to more than 4,000 apps. The company envisions "teams of dots working together on your behalf."
What is Instinct?
Instinct is an invite-only personal assistant that is gaining rapid adoption. The service does not require an app and is instead activated via text messages sent from your phone or computer.
Why Proof of Human matters to Muse, dots, and Instinct
Proof of Human is a way for someone to verify that they are a real, unique person without revealing who they are. An individual can extend their Proof of Human to their personal agents as they complete tasks via websites and apps that are designed for humans. This new paradigm has already produced friction. Shortly after launch, Amazon blocked Muse from shopping on its website, noting, among other objections, that the agent did not disclose that it was an AI agent while browsing, as Amazon's rules require.
As agents become more prevalent, websites and services will need to know what they are dealing with. That requires attribution. First, agents need to indicate where they are coming from and, ideally, who built them. The next step is establishing whether a real person stands behind the agent, which is where Proof of Human comes in.
Proof of Human is critical for a simple reason: The number of people on the planet is finite, roughly eight billion, while the number of AI agents is unlimited. Counting the people behind those agents matters for rate limiting (the number of queries an agent can make on behalf of a person), sales (the number of people interested in a product), and allocations (the number of products each person can buy).
How World ID helps
World ID lets a person present a variety of Proofs to third parties about themselves, including Proof of Human. Individuals can also delegate their Proofs to their agents, resulting in human-backed verified AI agents.
Proof of Human is stored in World ID app and uses zero-knowledge proofs to maintain privacy. When a Proof is delegated to an agent, websites and services can then verify that agent is acting on behalf of a person and set limits accordingly. Verified agents help increase trust and security for daily interactions.
Verified agents are:
- Private: No personal information is shared when a person or their agent presents a Proof.
- Unique: The Proof belongs to a person, and each person is counted once. Someone can control thousands of agents and still be counted as one human.
- Neutral: World ID is built on an open protocol, so it can work with agents created by any party, for any verified human, on any platform.
Verified agents in the real world
World is building out use cases that show why Proof of Human is important. One involves limited product releases, where each purchase must come from an agent representing a unique person: one human, one agent, one allocation. Another is human in the loop, where a person is prompted to approve an action before the agent completes it. If your AI agent wants to make a sizeable purchase or execute an agreement, you can verify that you approved it.
Building with World ID for Agents
Organizations are putting these tools to work. Okta has an early access beta of Human Principal, a service that can use World ID to verify that a real human stands behind an AI agent. Vercel is integrating World's human in the loop into its Workflow SDK, which lets developers add a World ID check to confirm a human approved a step in a workflow. Exa provided 100 free API requests per month to verified humans, shared across all the agents they back. Browserbase plans to use World ID so that human-verified agents can navigate websites with fewer blocks. These use cases were demonstrated at Lift Off.
What’s next
As agents proliferate and superintelligence advances, more websites and services will ask the same question: is a real person behind this request? Agents that cannot answer are likely to be blocked, which creates friction and frustration for the person who sent them. Proof of Human gives them an answer, and World's AgentKit, currently in beta, offers tools to build that answer into agents.
Visit world.org/world-id to learn more about World ID or docs.world.org to start building with World.
Human Intelligence is Surprising
Human intelligence is a recent evolutionary accident, making the difficult task of defining precise, unambiguous objectives for AI our most critical responsibility.
Summary
Deep Dive
- Frames evolution as a 'lazy' optimizer that produced general intelligence only as a side effect of survival.
- Contrasts evolutionary development (billions of years) with AI training (months).
- Highlights that recursive self-improvement processes (finding failures, data collection, model serving) are becoming procedural.
- Argues that benchmarks act as a worldview, where poorly defined reward functions lead to brittle or manipulative agent behavior.
- Emphasizes that our role is to act as architects of the 'hills' worth climbing.
Original Article
Human Intelligence is Surprising
It is common to be humbled by AI. Today’s models beat experts who have mastered their craft over years, often in minutes, and with simple prompts. As models continue to advance quickly, this is happening repeatedly and everywhere. Software engineers are watching the nature of their work change while writers, researchers, lawyers and bankers all stand nervously by.
Activities once reserved for truly gifted people can now be reproduced for pennies on the dollar. It’s natural that encounters like this lead some to more fundamental questions: what was it that made me special, what makes humans special?
Looking for some remote edges of human capability is the wrong response; we should take an even wider and more humbling view. Examining the origins of intelligence reveals that it is not the inevitable endpoint of evolution, but a relatively recent, circumstantial, and surprising phenomenon.
Life appeared on Earth at least 3.7 billion years ago. Complex multicellular organisms emerged roughly a billion years ago, animals around 600 million years ago, land plants around 470 million years ago and the first primates around 55 million years ago. Homo sapiens have existed for only about 300,000 years. On the timescale of life, human intelligence arrived almost yesterday.
And once it arrived, it moved quickly. A species with enough intelligence for language and toolmaking learned to reshape its environment, leading it to exponential and cascading planetary change. What feels to us like the long arc of human history and complex cultural development is, from the perspective of evolution, a sudden discontinuity.
In fact, we actually never really had much of a chance to develop meaningful intelligence. Evolution does not optimize for intelligence in the abstract and has no implicit reason to produce the maximum possible intelligence. In effect, once we saw a glimpse of consciousness, we began changing our environment faster than natural selection could change us. We had crossed the threshold at which our technology guaranteed our existence.
Evolution is a lazy optimizer for intelligence.
It should not surprise us that there are better mechanisms for developing particular forms of intelligence. Evolution takes billions of years, haphazardly optimizing organisms for survival under messy and shifting constraints. Machine learning can optimize directly for a target that we explicitly define.
Consider a definition of intelligence as the ability to master the game of Go.
Humans were obviously not selected by evolution for their ability to play this game. The fact that a person can study Go, understand its strategy, and eventually become a grandmaster is a fortunate byproduct of a general intelligence that evolved for entirely different reasons (survive and reproduce in their particular environment). Yet once Go performance became the explicit objective, researchers at Google built AlphaGo, which quickly became superior to the best human Go player at the time.
The achievement took only on the order of months once development focused on that objective. It isn’t because Go is trivial. It is because the optimization target was made unambiguous. Once the hill was visible, we were able to climb it quickly.
When researchers can specify a task, construct a useful distribution of problems, and measure success, AI can reach superior-to-human performance in an extremely short period of time. Not every important capability is easy to formalize. But I expect the hill-climbing trend to continue wherever we can explicitly capture a task that “smart” humans can do today.
Defining intelligence is the greatest barrier to realizing it.
The emerging work on recursive self-improvement (RSI) makes this especially clear. An observer following the lab news of crazy compute spend, outrageous compensation packages and magical capabilities of models may believe the model development process to be mystical. In actuality, the model development process is now well understood and somewhat procedural. Model development can even be decomposed into tasks: finding failure modes, collecting data, running experiments, optimizing the infrastructure, and serving models, to name a few.
In our effort to define and measure RSI through a set of illustrative small tasks from this process, we now have well-defined metrics we can contribute to our conception of intelligence itself. Hill-climbing as a dimension of intelligence is made legible. And sure enough, we are starting to see AI models show promise at the discrete steps for RSI. As with all prior hills, hill-climbing is itself hill-climbable.
This reinforces my belief that being able to define what intelligence should look like is one of the last real human problems. Even the hill-climbing process itself will soon be automated.
The optimization process is becoming well defined and therefore easier to optimize.
This is discerning, but it also unveils a more foundational and disconcerting problem. If defining intelligence allows us to build it, then the definition matters enormously. A system optimized to earn money in the stock market will develop a different intelligence from one optimized to cure cancer, write requested software, or play Pokémon. Even within a single domain, a slightly wrong objective can produce behavior that is highly capable and completely undesirable. A coding agent rewarded only for passing a test might fix the software or manipulate the test, exploit the environment, or take some shortcut its designers failed to anticipate.
This is not a new problem introduced by AI. Human intelligence has always developed in response to how it is evaluated. Students study reading, writing, and mathematics because those subjects are tested in the SAT and determine college placements. In eras when Latin signaled education and status, students spent years learning Latin. Labor markets influence where ambitious people spend their time improving and applying their abilities: trading on the market, trying to solve cancer, writing a novel, or elsewhere.
Measures of intelligence shape it. Defining intelligence is our greatest barrier to realizing it.
With Vals AI, we care a lot about exposing model capabilities through well defined evaluations. Our primary tools for this are the benchmarks we construct measuring models’ ability to do real, valuable work.
In my most abstraction conception, a benchmark is an input distribution and an output-verifier. These are the problems or queries we choose to present, and the way we decide what counts as a good answer, respectively. Both choices carry a worldview. If the problem distribution is narrow or contrived, a system can become excellent at the test while remaining brittle outside it. If the reward captures the wrong thing, the system can become excellent in exactly the wrong direction. It is the burden of the person making the evaluations to ensure they are pushing model progress in the right direction.
Optimizing for the wrong or an ill-defined notion of intelligence will lead us astray.
The rise of machine intelligence should make us doubly humble. First, we should be humble about human intelligence. It is extraordinary, but it is also recent, circumstantial, and shaped by an optimizer that was never trying to create the smartest possible being. There is no reason to assume it represents an upper bound.
Second, we should be humble about our ability to define what comes next. As optimization becomes cheaper and more powerful, choosing the objective becomes more consequential. The most important remaining human task may not be to stay ahead on every hill but instead to decide which hills are worth climbing and how to construct them. Whereas for humans intelligence has largely plateaued, for models there is always a higher peak.
Meta open sources code to let you make Muse AI gadgets
Meta is releasing open-source code enabling developers to integrate their Muse AI agent into custom hardware projects like E Ink displays and Raspberry Pis.
Summary
Decoder
- ESP32: A low-cost, low-power system-on-a-chip microcontroller with integrated Wi-Fi and dual-mode Bluetooth commonly used for IoT and DIY hardware projects.
Original Article
Full article content is not available for inline reading.
Apple's John Ternus Has Found a New Design Chief: Himself
Apple executive John Ternus is taking direct oversight of hardware and software design, opting not to replace the former Chief Design Officer.
Summary
Original Article
Apple CEO John Ternus is spending a lot of time overseeing the company's hardware and software design work. He is largely leaving the company's financial operations to CFO Kevan Parekh and the supply chain to COO Sabih Khan. Ternus is deeply involved in product engineering and design. One of his first priorities was to restore the company's design studio standing. The company is not looking to recruit an outside chief design officer to oversee the entire organization.
Superpowers Race to Put Nuclear Reactors on the Moon
NASA is soliciting contractors to develop a low-maintenance lunar nuclear reactor by 2030 to secure energy for long-term moon base operations.
Summary
Original Article
NASA wants contractors to prepare to build a nuclear reactor that could survive a space voyage and run without maintenance near the south pole by December 2030. The agency is racing for the future of space exploration. The US and China are both racing to establish a moon base to search for resources and to launch missions deeper into the solar system. Controlling access to the Moon's assets is vital to this mission. Both superpowers believe a nuclear reactor is the centerpiece of this ambition.
I Quit OpenAI Because Its Culture Is Broken
David Robinson resigned from OpenAI’s safety team, citing the company's reckless culture of iteration as increasingly untenable for building safe, high-stakes AI.
Summary
Original Article
David Robinson resigned from OpenAI this week, joining a parade of former colleagues who have decided that the company's current path is unacceptable. Robinson headed the writing of the safety reports that OpenAI publishes with each major launch. OpenAI has thrived by trial and error, but this approach guarantees periodic failures, and the scale of those failures is growing. Achieving something much closer to perfect the first time is becoming more essential because iteration after a mistake may not be possible.
The Mulleted, Meme-Loving Billionaire Behind Meta's Hit AI App
Meta is betting on 29-year-old executive Alexandr Wang to bridge the gap with younger users through meme-driven marketing and a Gen Z-native approach.
Summary
Original Article
Alexandr Wang, 29, was hand-picked by Mark Zuckerberg last year to oversee Meta's AI efforts. Wang, the company's first Gen Z senior executive, brings an internet-flavored communication style that could help the company connect with younger customers. He is a native of the social-media ecosystem his boss built. Wang has succeeded in building hype for Muse, in part with an unconventional internet campaign. The challenge is now to keep the momentum going as rival companies press ahead with their own offerings.
The Sleuths Who Expose When AI Goes Rogue
Online sleuths have formed decentralized groups to track and document 'swarm' incidents where AI agents interact in unexpected, rogue ways.
Summary
Decoder
- Swarm Incident: A situation where multiple autonomous AI agents interact or influence each other in ways that lead to unpredictable or emergent behavior.
Original Article
The swarm chasers are a loose crew of online sleuths that work together to piece together interconnected AI swarm incidents.
So You Think You Could Be An Electrician?
The narrative that AI-proof blue-collar trades are an easy alternative to college degrees obscures the brutal physical reality and internal hierarchies of those professions.
Summary
Original Article
Funny how the people telling kids to get jobs in the trades tend to be white-collar.
There’s a popular story about working in the trades that goes something like this: College prices have spiraled wildly out of control, while trades salaries have skyrocketed, often surpassing those of college graduates. Even if you do get a degree, AI will displace most knowledge work. So why would anyone, the story goes, spend $200,000 on a worthless college education when union electricians make six figures performing work that’s impossible to replace with AI? There’s a populist angle here, too: The nerds with guaranteed B averages who waltzed into high-paying jobs will get their comeuppance as they’re pushed aside by the hard-working blue-collar moderates who ceaselessly toil at the real work of building the country!
The skilled trades hold a special place in the public imagination as a route out of poverty. Trades have advocates like Mike Rowe, former opera singer and host of the TV show Dirty Jobs, whose foundation routes smart, hard-working kids into the trades via scholarships. (Today, its website prominently advertises “AI-proof six-figure jobs.”) But Rowe has something in common with most people selling this popular narrative: a college degree.
Trades workers tell a very different story. When I was an apprentice over 30 years ago, we would tease the old-timers for badgering us to quit construction and return to school (the “look at these hands” lecture). In fact, one of the most reliable indicators of an apprentice’s low intelligence was being quickly embraced by old-timers, because it meant they immediately recognized that the kid’s greatest potential was in construction.
No one understood the physical demands and uncertain rewards of trades jobs better than older tradespeople. So take it from an older tradesman: The story is more complicated than you might think.
Less money
In the first place, college graduates still earn more than those without college degrees, specifically because of their academic credentials. In casual settings, people often fail to disentangle selection effects, but rigorous causal research actually does find substantial income effects of college graduation, often amounting to hundreds of thousands of dollars over the course of a career, though this can vary across majors. Completion of a four-year degree is critical to the earnings leap.
It is also not the case that wages in the trades are rapidly increasing. The college wage premium grew steadily for several decades beginning in the 1970s until the 2000s. There was a brief reversal period around COVID, during which most gains went to a narrow range of subtrades. This is hardly sufficient to warrant an endorsement of trades employment based on income.
There are, of course, also potential missteps. Social work majors on average earn much less than their college peers, and the degree places students on a high-probability collision course with a low-yield graduate degree. Nevertheless, for decades the staple advice that high school graduates should attend college has probably been sound. In this regard, the old-timers with calloused hands who pressed me to leave construction were probably right.
This holds particularly true when considering the real cost of college. While the median sticker price for four-year college has dramatically increased over the past 30 years, most students pay much less than the advertised price. The average net paid price for public in-state college tuition and fees was under $2,500 per year in 2024 and has been declining since 2016.
So: The actual price paid for college isn’t very high and is declining. Obtaining a college degree unlocks higher wages. Plus, white-collar jobs are safe and easy.
More physical harm
When I was 20, I broke both my heels in a construction accident. The right one was surgically reassembled using four heel fragments and extra bone from my hip. Workplace accidents seem to have worse outcomes than similar accidents occurring elsewhere, and at the time this specific injury had a generally poor prognosis, so doctors were initially uncertain whether I’d be permanently disabled.
When I think about my teenage son or daughter working in construction, my mind often goes to The Accident. I picture my daughter, who’s three years younger than I was at the time, wheeling to The Beer Store in a Workers’ Comp rental wheelchair each day, and coming home to down 12-pack of Molson Canadian. Or I think of my son sleeping with a fan every night to overpower the ringing in his ears. Over enough time, physical jobs exact a physical toll that desk jobs largely avoid.
Most white-collar professions have a trivially small risk of death. Even in gray-collar jobs that are widely perceived as dangerous, such as policing, most workplace deaths are caused by motor vehicles. But some trades jobs carry a risk of death or serious injury connected to specific and inherent job demands. That is, something that seems dangerous that you often do is the most likely to harm or kill you.
Jobs performed at heights, such as ironwork or roofing, have a high mortality due to falling. Electrical linesmen tend to be killed by electrocution. Arborists can be killed in a variety of ways: falling from heights, being struck by trees or electrocuted. For people in these jobs, these risks are widely understood and accepted. Almost every tradesperson knows someone who’s had at least one protracted absence due to a work injury, and many of us know someone who was permanently disabled or killed. These risks are probably less salient for people who have worked mainly white-collar jobs and know few people employed in dangerous professions.
There’s moderate evidence that people in wealthy countries are increasingly risk-averse to things that might kill or harm them. For many folks, this may translate to an unwillingness to take on a nontrivial risk of death or injury via dangerous jobs. And there may be a subtler version of this in which people are reluctant to tinker with physical objects for fear of destroying them. I’ve found that when I coach people through DIY projects, they’re often extraordinarily fearful. And not just of electrocution or falling, but of doing something incorrectly or pushing a broken thing past the point of being reparable. (And this appears to be something that blue-collar people — perhaps due to hubris or ignorance — don’t seem to worry about as much.)
While the optimal quantity of risk-taking for a given job probably isn’t zero, it’s difficult to broadly argue against placing a high intrinsic value on human life or the avoidance of pain and chronic disability. People abandoning the trades for safe, high-paying jobs is understandable.
Internal status hierarchies
White-collar people often seem attuned to the differences in social status and compensation of white-collar jobs while remaining somewhat oblivious to the same internal dynamics across blue-collar jobs. Most realize that a partner at McKinsey is very different from a bank teller and that the great majority of college graduates have essentially no path toward McKinsey partnership.
But they’re often incapable of generalizing this to understanding the trades, or contrasting this against the backdrop of blue-collar employment. The median blue-collar worker isn’t in the trades at all: They’re probably a delivery driver for Amazon making roughly $40,000 per year. If lucrative trades employment is as simple as shaking the union electrician tree until a $125,000-per-year job falls out, why don’t delivery drivers simply pivot toward becoming one? The answer is that within the blue-collar status hierarchy, the union electrician is much closer to the McKinsey partner than the bank teller.
Trades jobs are generally among the best blue-collar jobs, and within the trades, specialized union jobs are especially coveted. These jobs are less accessible to people on the lower end of the blue-collar hierarchy. They often require training, social connections, certifications and licensing, and technical competence. I asked several acquaintances in union trades about the selection process for their jobs. Each described an initial process with 100 to 200 candidates, followed by testing and interviews that winnowed the field to one to two dozen. The final selection of between four and eight job entrants was less clear but appeared to include a mix of connections, previous credentials such as military service, and measured ability.
Public schools often have standardized tests and exams that allow for objective comparisons between students, but in the U.S., trade schools lack uniform performance metrics. Consequently, the effectiveness of any trade school is hard to determine. But we should be wary of adopting the widespread assumption that training centers — which, for have decades, have acted as a repository for young people who are unable, for academic reasons, to attend college — have better teaching methods than other schools.
AGI won’t spare the builders
What if the arrival of artificial general intelligence eliminates or substantially reduces demand for knowledge workers? Could this reshuffle the employment deck in favor of trades?
I won’t claim to know for sure: The direction and degree of AGI’s impact on employment are still unclear and hotly contested by people much smarter than I. As of this writing, large language models are much closer to human-level performance on white-collar metrics than robots are on the work of tradespeople. However, one underdiscussed part of this conversation is how AI will affect the trades themselves.
People often understand successful innovations that are service-focused (Airbnb), pharmaceutical breakthroughs (mRNA vaccines), or even radically new physical technologies (nuclear power), but don’t understand the iterative nature of technological changes to physical domains like manufacturing and construction. The trajectory of innovation in the trades is away from physical dexterity. Today, innovations often bypass dexterity instead of improving or refining it. This characteristic is common to process improvements in physical domains, but it isn’t intuitive for people familiar with the ground-breaking innovations of the past few decades.
Framing job sites of the 1950s and ’60s had many carpenters who could drive 16 penny (3.5”) nails with two rapid blows — the first to slightly set the nail and the second to fully drive it. Someone evaluating the process of mechanical fastening at the time could have reasonably assumed that making improvements to hammers, one of the few tools that predates Homo sapiens, was a fool’s errand.
But by the 1990s, pneumatic nailers — trigger-operated tools that used compressed air to drive nails with a single bump or trigger press — had largely replaced hand nailing on framing jobs, reducing demand for hammer proficiency and speeding production. Pneumatic nailing entirely bypassed hammers and the physical act of driving nails. We can see the same dynamic with the transition to rebar tie tools and auto-fed screw guns. While it’s difficult to project the specific innovations a superintelligent machine might bring to the construction industry, it seems likely that some will follow this trend.
And if some follow this trend, we might not need robotics to advance all the way to human-level handling of current tools before we see significant reductions in the need for human labor.
As dexterity demands have decreased, work in the trades has also gotten more cognitively demanding. It’s likely that AI could reverse this.
This increasing complexity is a result of the broad push toward mechanization and away from manual inputs that are common to other fields such as agriculture and manufacturing. At the same time, buildings have become more complicated partly in response to increasingly unique consumer preferences and stringent building codes. Residential framing today is often aided by telehandlers, which function as mini-cranes that can lift whole building assemblies into place. Some systems within buildings, such as HVAC, are essentially governed by complex, installer-programmed computing systems. Furthermore, these systems are increasingly integrated across platforms. For example, smoke and carbon monoxide alarms might be inputs for heating and cooling systems to reduce fire spread. Optimizing these systems currently requires astute and well-trained technicians to stage and monitor them.
With improvements to video interpretation capabilities, AI could remove much of the complexity from the minds of human tradespeople. For instance, a future job site worker equipped with smart glasses and an earpiece might receive continuous direction via an AI interpreting and providing feedback on their actions. In some cases this is similar to the way I occasionally use LLMs, although at a lower level. It’s much faster to ask for torque specifications via an earpiece than to reference a manual. One of the most failure-prone jobs in residential HVAC, for example, is setting airflow. A technician might install or service a dozen manufacturers, each with three varieties of fan type and dozens of models each with unique switch set points. On complex HVAC installations there may be dozens or even hundreds of reference points such as these. Many are nearly impossible to memorize or reasonably guess. An AI with good access to the technician’s field of view and perfect equipment knowledge could readily coach dumb hands through this task.
Advice time
How can we provide better counsel to people considering a career in the trades? First, we should be as frank about the negatives as we are about the positives. We should also pay attention to who we’re talking to. The best advice will depend on whether our audience is a high school student from a blue-collar family looking to claw her way into the middle class, or a software engineer concerned that his job will soon be replaced by a recursively self-improving artificial intelligence.
The trades are more physical than working in an office, and some people find this — and working outdoors, where much trades work is performed — liberating. Some derive great pleasure from achieving a flow-like state of mind from performing specific physical tasks. Some people find it deeply rewarding to physically make things and to see progress and improvement in doing so over time. Some people derive a deep sense of camaraderie from physically toiling with other people. It could be that any of these, especially when combined with a low sensitivity toward financial compensation or status, makes the trades a better fit than white-collar employment.
In a world where white-collar people are diverted from traditional office jobs to blue-collar ones, they would likely have strengths and weaknesses relative to the existing labor force. Someone with a strong background in programming might quickly be better than many existing workers at HVAC controls. On the other hand, adapting to highly physically demanding jobs would be more difficult, so roofing or formwork carpentry would have a longer adjustment period. We might also want to match candidates based on traits like introversion and extroversion. For instance, some trades traditionally work largely in isolation and are culturally sedate, such as woodworking. Others have near continuous interaction and tend to be culturally boisterous, such as framing.
There’s one thing we need to avoid: reducing the trades to a simple morality tale. A trade job can be a great fit for the right person, but it’s not a magic ticket to prosperity or a guaranteed hedge against AI automation. They’re also not a monolith. If you’re interested in training or retraining as a roofer or a plumber, an electrician or an HVAC installer, don’t waste your time listening to opinions on “the trades” — talk to the people who do them.
What Meta Got Right With Muse
Meta’s Muse model succeeds not through technological breakthroughs, but through a superior combination of existing components.
Summary
Original Article
None of Muse's capabilities are new, Meta just put the pieces together in the right way.
Spotify billionaire's body scan startup has come to America
Neko Health, founded by Spotify's Daniel Ek and Hjalmar Nilsonne, is expanding to the U.S. to offer preventative body scans for $500.
Summary
Deep Dive
- Neko Health uses computer vision and blood analysis to screen patients.
- The service requires a $500 out-of-pocket payment per scan.
- Currently operating in New York with plans for national expansion.
- The company claims a global waitlist of over 300,000 individuals.
- Investors emphasize the role of AI in scaling preventative medicine.
Decoder
- HSA (Health Savings Account): A tax-advantaged savings account in the U.S. for people who have a high-deductible health insurance plan, allowing funds to be used for qualified medical expenses.
- Biomarker: A measurable indicator of some biological state or condition, in this case, blood-based metrics used to assess health status.
Original Article
Neko Health seems to be the hottest body scan in the world right now.
The scan, co-founded by Spotify’s David Ek and Hjalmar Nilsonne, uses advanced technology to screen for issues in the blood, skin, and metabolism. The company was founded in Sweden, has raised nearly $1 billion total, and just landed in New York. It has a global waitlist already topping 300,000 people, the company says.
TechCrunch sat down with Farooq Abbasi of Preface Ventures, who invested in the company. He hopes Neko Health can be the future of preventative medicine, especially in a country like the United States, where the medical system is notoriously plagued with issues.
Abbasi contends that preventative care utilization is low in the U.S. as people have concerns about how much it costs. Currently, Neko Health charges $500 for its one-hour scan, which is not covered by insurance but should qualify as an acceptable expense for those that have health savings accounts (HSAs). “I think preventative care and healthcare should be cheap and accessible,” Abbasi said.
The hope is that Neko Health can tackle some of these issues, and it’s currently eyeing an expansion across the U.S. Abbasi has been scanned twice already and, perhaps, no surprise, says he loves it. “I think there should be Neko clinics everywhere,” he said.
He says the product has grown so much since he first tried it. “In my generation one scan, it was a blood test with 15 biomarkers,” he said. “It was fifty-five this time.”
Neko also uses advanced computer vision for its scans and Abbasi sees this as another example of how AI is improving the medical sector overall. He argues that Neko Health isn’t just another wellness fad, but something that gives people “actionable insights to make their lives better,” he said.
Neko still has a long way to go in the U.S., especially on scaling and regulatory approval. But Abbasi is an optimist; he hopes more venture dollars will pour into the right companies and doesn’t think anything is too broken to fix.
CD Foundation Project Updates September 2026
The CD Foundation's September update features major Spinnaker infrastructure cleanup and Jenkins hardening efforts.
Summary
Decoder
- Halyard: A legacy configuration and management tool for Spinnaker environments, now replaced by kustomize.
Original Article
The CD Foundation currently has 6 open source projects: CDEvents, Jenkins, Jenkins X, Ortelius, Screwdriver, and Spinnaker.
Our projects solve some of the biggest issues in the Continuous Delivery and Continuous Integration space. We always want to expand our project communities with additional contributors, end users, and passionate technologists. Please check out each project to see what you can do to get involved!
Here are the mid-year highlights for each project Features and Releases.
CDEvents
Features and releases
The project also started modelling the concepts used in links as schemas, so that SDKs and users can share and validate them:
- Links can now reference a domain-level identifier (
domainId), a CDEvent context ID, or both. This change also introduced a governed registry of providers. - Spinnaker is being added to the provider registry.
- A registry of resource types for the
typesegment ofdomainId.
Other changes merged in the spec for v0.6:
- A new
queuedpredicate fortaskRun, contributed by a new contributor. - A consistency fix for the
taskRun.finishedevent.
New proposals under discussion include changes to the deployment events, subjects and events for scheduled executions, change windows, data migrations and reconciliation, an optional canonical hash for event integrity, and event types for feature flags and SCM branch pushes. The project communnity is aiming for a v0.6 release after these proposals are reviewed.
Releases:
- Spec v0.5.1 (April 2026)
- Rust SDK v0.4.0 and v0.4.1, with support for spec v0.5.1
- Go SDK v0.5.1, with support for spec v0.5.1 and bug fixes
Security Updates
- The Rust, Go and Java SDKs added a Plumber security scan to CI and hardened their GitHub Actions workflows.
- The project communnity is discussing event-level provenance and tamper evidence, in light of the verifiability extension recently merged in CloudEvents.
JayeX
Gateway API support
The JayeX applications has gotten support for Gateway API as an alternative to Ingress. As a default gateway implementation Envoy Gateway has been choosen and baseline configuration has been added. More details can be found in the issue.
Upgrade rally
Our go applications and libraries (we have many) has been upgraded to go 1.24 and also a lot of dependencies have been upgraded.
Project name change
Work has been ongoing to switch the project name to JayeX. So far the web site and the plugins are done.
Jenkins
Google Summer of Code 2026
Jenkins has completed five Google Summer of Code 2026 projects. Read more about three of the projects in their concluding blog posts:
- Retooling Jenkins.io Web Success Stories by Vatsal Verma
- Jenkins Email Notifications using Outlook SMTP with OAuth by Mohammed Faheem
- Plugin Modernizer Stats Visualization by Pratik Mane
Special thanks to our organization lead mentor, Kris Stern, and the other Google Summer of Code mentors.
Features and releases
Notable weekly releases in Q2 and Q3 2026 included:
- Jenkins 2.580 (September 2) – Security fixes as published in the security advisory
- Jenkins 2.579 (August 11) – Apache Commons Lang 2.6 removed from Jenkins core
- Jenkins 2.577 (August 11) – Jenkins binary is over 40 MB smaller than it was a year ago
- Jenkins 2.574 (July 21) – Jenkins binary is over 20 MB smaller than it was a year ago
- Jenkins 2.568 (June 10) – Security fixes as published in the security advisory
- Jenkins 2.562 (April 28) – Jenkins MSI installer is now signed with a certificate provided by the Linux Foundation
LTS releases in Q2 and Q3 2026:
- Jenkins 2.568.3 (September 2) – Security fixes as published in the security advisory
- Jenkins 2.568.2 (August 5) – Security fixes as published in the security advisory
- Jenkins 2.555.3 (June 10) – Security fixes as published in the security advisory
- Jenkins 2.555.2 (May 13) – Jenkins MSI installer is now signed with a certificate provided by the Linux Foundation
Jenkins continued its pattern of releasing a new version every week and a new long term support version every 4 weeks. New features in long term support releases are summarized in the “What’s new in Jenkins LTS” playlist.
Security Updates
The Jenkins security team continues to track security issues, report vulnerabilities, and resolve security issues.
Security advisories were published in Q2 and Q3 2026 included:
- 29 April 2026 Plugins Security Advisory
- 27 May 2026 Plugins Security Advisory
- 10 June 2026 Core Security Advisory
- 24 June 2026 Plugins Security Advisory
- 5 August 2026 Core and Plugins Security Advisory
- 2 September 2026 Core and Plugins Security Advisory
Ortelius
Roadmap
- Continue to refine the Ortelius AI for updating package manager files using the Claude LLM and Ortelius MCP.
- Next is to add in GitLab support for getting the release and deployment data from GitLab workflow logs.
- Add more “how to” videos.
- Onboarding -delivering value in under 10 minutes.
Features and release
- Last big release in Q1 – V12 release for the new UI and backend using pure ArangoDB. REST, GraphQL and Kafka are now supported. This also includes easy onboarding using a GitHub Ortelius App to bring in release artifacts and deployment data. The onboarding continues to be worked on.
Adoption updates
We continue to work toward a corporate end-user. Work has been done to simplify the onboarding with auto scanning. The Ortelius Outreach has created a “Pathfinder” badge to encourage onboarding open-source projects. OS projects only have repos, but not endpoints. Ortelius onboarding will show vulns in the releases, not the entire repo, based on SBOMs.
Security updates
Ortelius is now monitoring itself. We continue to use revonate on all (~40) repos keeping dependencies up to date. We are now focused on those CVEs that are being released and those are fixed within 2-3 days.
Screwdriver
Features and release
-
Component versions as of Q3 2026:
- API v8.0.192
- UI v1.0.1420
- Store v7.0.4
- Queue-Service v5.0.11
- Launcher v6.0.237
- Build Cluster Worker v6.1.2
Notable Updates:
- Support for partial versions when using templates.
- Expanded branch name character set
- Ability to disable a pipeline
- Lots of updates and improvements to the UI/UX
Contribution update
We are pleased to have raihata join as our newest contributor to the project. Project continues to maintain a steady pace of development.
Spinnaker
- 2026.3.0 is released with MAJOR improvements and enhancements. Highlights:
- AWS SDK1 is removed
- New UI’s for account management
- SAML moved to spring config
- Deprecation announcemnts on multiple tools/configs/resources
- MCP Server
- Halyard was removed from the codebase and kustomize is now the documented and supported installation method
- Post release, cleanup of the old spinnaker images is about to be applied to purge data. This should reduce project costs, and keep newer supported images. GHCR being now the default, GAR will no longer be supported and images will be removed there after a period of time.
Features and release
Currently supported releases and release notes are available here
Contribution trends
The number of pull requests has drastically increased in the last two months. This is driven a lot by AI improvements to the project and major recent changes.
We are seeing some new contributions and it’s worth reviewing the stats.
Security updates
There are critical and high vulnerabilities, the most recent of which was published early July. It’s highly recommended to be on the most recent minor version of a supported release (the most 3 recent releases).
Infrastructure updates
Automatic image cleanup. This will be purging old release images – this will be cleaning up almost 10 years of old images as a result.
For more information on the CD Foundation projects and how to contribute, check out the project page.
Diving Through Data at OpenAI: How a Data Agent Navigates 70,000 Datasets (27 minute video)
OpenAI successfully reduced reliance on model raw intelligence by building an internal data agent that prioritizes schema lineage and continuous learning.
Summary
Original Article
OpenAI built an internal AI data agent that answers complex data questions by combining warehouse queries with rich context from schemas, lineage, code, company knowledge, memory, and live sources, while showing its assumptions, SQL, citations, and confidence. The big lesson is that reliable data agents depend less on raw model intelligence and more on giving them the right context, continuously learning from corrections, and testing them with strong evals.
Microsoft, Google back Apache Ossie to make enterprise data and AI platforms more interoperable
Microsoft and Google are backing Apache Ossie to standardize semantic models, attempting to reduce vendor lock-in for enterprise data platforms.
Summary
Decoder
- Semantic model: A representation of business data that defines relationships, metrics, and terminology, ensuring consistent logic across reports and applications.
- Metric drift: The phenomenon where the same business calculation yields different results because it is defined differently across multiple analytics tools.
Original Article
Microsoft and Google are joining a project to create an open specification for exchanging semantic models across data, analytics, and AI platforms. The project already has the backing of over 60 companies including Databricks, Informatica, Mistral AI, Nvidia, Oracle, Salesforce, and Snowflake.
Their support for Ossie makes it a little more likely that future analytics platforms will be interoperable, but is no guarantee that vendor lock-in will go away, analysts said.
The project began life as Open Semantic Interchange (OSI), then became Apache Ossie when it was accepted into the Apache Incubator in June. It uses JSON and YAML to represent semantic models, including datasets, fields, relationships, metrics and AI context, making them interoperable across platforms such as Snowflake, Databricks, Tableau, ThoughtSpot and Sigma.
Ossie uses a hub-and-spoke model for interoperability, serving as the common semantic format while converters translate between it and individual vendors’ semantic implementations. This means an enterprise could move a semantic model between compatible platforms without requiring a separate point-to-point converter for every pair of vendors.
Microsoft’s support for Ossie includes developing a two-way converter between Power BI semantic models and Ossie, allowing semantic context defined in Power BI to be represented in the common format and Ossie models to be converted back into Power BI.
It is also pushing for the Ossie specification to include more support for its ontologies, and for Ossie’s recognized query languages to include DAX, the expression language used by Power BI for defining calculations and business metrics.
That means enterprises could carry the calculation logic that gives their semantic models business meaning alongside the underlying model when moving between platforms that support Ossie.
Google, meanwhile, is in the process of joining Apache Ossie, a company representative said via email response. The company has not yet provided details on what it plans to contribute or has already contributed.
The Ossie specification lists BigQuery/GoogleSQL as a supported dialect, a result of “early community contributions that recognize BigQuery’s footprint across enterprise data stacks,” the representative said.
Microsoft, Google backing could reduce engineering work
Support for the project from the two technology giants could reduce the burden of repeated semantic engineering for enterprise teams moving analytics workloads between platforms, said Aditya Ranjan, senior data engineer at supermarket giant H-E-B.
That reduced engineering overhead could also help avoid metric drift, which is often the consequence of recreating and maintaining the same business definitions across platforms, according to Ashish Chaturvedi, executive research leader at HFS Research. With Ossie, developers can instead define a metric once and treat it as a versioned, reviewable code artifact, rather than rebuilding and maintaining it separately for each platform, he said.
Another potential benefit, said Ranjan, is that productivity could improve if developers spend less time translating and validating semantic definitions for each platform when onboarding new analytics tools or deploying new analytics or AI applications.
It can also make it easier for developers to build agentic applications with consistent business context, since maintaining one definition of a metric across platforms could reduce the risk of different agents interpreting the same metric differently, said Stephanie Walter, practice leader of AI stack at HyperFrame Research. That could give CIOs greater confidence to scale agentic deployments, she added.
Portability
For Michael Leone, principal analyst at Moor Strategy and Insights, the bigger advantage of Microsoft and Google supporting Ossie is that it gives enterprises more ownership of their semantic models by making them more portable.
That portability gives buyers more leverage, said Walter, since changing a BI or data platform would no longer require rebuilding the semantic layer from scratch.
However, portability does not necessarily mean complete interoperability, as the degree to which enterprises can move semantic models between platforms will depend on how much vendor-specific logic and functionality Ossie can represent and how accurately the destination platform can interpret it, Walter said.
“A portable structure does not guarantee equivalent behavior. A metric may survive conversion syntactically but still produce a different result because the destination interprets joins, nulls, time calculations, or filters differently.”
That challenge is particularly relevant for Power BI, as its time-intelligence functions, calculation groups, and context transitions in DAX do not always map cleanly to ANSI SQL, meaning a sophisticated Power BI model could lose some of its behavior in translation, Chaturvedi said.
CIOs, therefore, will still need to validate whether converted models preserve the intended calculations, business logic and results before putting them into production, especially for complex ones, Ranjan advised.
There are gaps around governance as well, Chaturvedi said: “The core specification covers datasets, relationships, fields, and metrics, but doesn’t mention row-level security, access policies, or certification status as first-class elements.”
That means enterprises have to set them up again on each platform.
Moreover, Ossie’s stability and long-term adoption remain open questions. The project is still “incubating”, meaning that it isn’t a full Apache Software Foundation project and has yet to prove it meets community norms. The current specification version, 0.2, is explicitly labeled a development draft, meaning the schema and capabilities could still change as it evolves, giving CIOs less certainty about its suitability as a long-term interoperability layer, Walter said.
“The current draft also removes the earlier document structure and does not yet define bundles or cross-model references, which are capabilities that large enterprises with interconnected models are likely to need,” said Chaturvedi.
Ossie could shift rather than eliminate vendor lock-in
Even if Ossie becomes widely adopted, it is unlikely to eliminate vendor lock-in entirely.
“New forms of lock-in are likely to emerge. Execution engines will retain proprietary behavior that no interchange format captures and vendor extensions will carry increasingly valuable features,” Chaturvedi said.
For enterprises, that means the definitions become portable while the behavior around them stays platform-specific, he added.
Apple's rumored smart home camera could take a different approach to solving privacy concerns
Apple is reportedly developing a privacy-focused smart home camera that processes footage locally and alerts users via text instead of storing video.
Summary
Original Article
Apple may be developing a privacy-focused smart home camera that uses AI-powered facial recognition to identify people and pets and understand what is happening around the home. Instead of recording video, the device could interpret activity locally and send text-based descriptions and alerts, without storing footage. The approach could differentiate Apple from conventional security cameras by making environmental awareness and privacy, rather than surveillance footage, the core experience.
Empathy Mapping: The First Step in Design Thinking
Empathy maps are collaborative tools for synthesizing user research, but they must be grounded in qualitative data to avoid becoming mere collections of assumptions.
Summary
Deep Dive
- Quadrant Structure: Categorizes user research into what users Say, Think, Do, and Feel.
- Research Grounding: Every sticky note must be traceable to a specific source like an interview or survey.
- Contradiction Analysis: Identifying conflicts (e.g., a user doing one thing but saying another) is the most valuable part of the map.
- Iterative Process: Maps should be updated as new research comes in; they are not one-time workshop artifacts.
- AI Usage: AI should only be used to organize existing data, not to fabricate user insights.
Decoder
- Empathy Map: A design tool used to visualize what users say, think, do, and feel to help team members build a shared understanding of user needs.
Original Article
Empathy Mapping: The First Step in Design Thinking
UX professionals advocate on users’ behalf. To do that effectively, they need to not only deeply understand users but also communicate this understanding clearly to colleagues and help prioritize users’ needs. Empathy maps, widely used across Agile and design teams, are a simple and durable tool for both of these goals.
Definition: Empathy Map
An empathy map is a collaborative visualization that captures what we know about a particular user or about a user segment. It externalizes knowledge about users to create a shared understanding of user needs and aid in decision making.
Empathy mapping can be driven by any qualitative research method, and a map can be sketched even if research is lacking, as long as the team is honest and clearly distinguishes evidence from assumptions. Maps help UX professionals understand what they know about their users and where they need to gather more data.
One User vs. Multiple-Users Empathy Maps
An empathy map may represent a single actual person (one-user/individual empathy map) or it can be an aggregate of data coming from multiple users (aggregated empathy maps) that belong to the same user segment.
Aggregated empathy maps represent a user segment rather than one particular user. They are usually created by combining individual empathy maps from users who exhibit similar behaviors and can be grouped into one segment. The aggregated empathy map synthesizes themes across that group and can be a first step toward creating personas. (However, empathy maps are not a replacement for personas, but they are one way to visualize what you know about a persona in an organized, empathetic form.)
Format of an Empathy Map
Traditional empathy maps are split into 4 quadrants — Says, Thinks, Does, and Feels — with the user or persona in the middle. The map shows the user as a whole and is deliberately not chronological or sequential. Whereas a journey map follows a sequence of steps, an empathy map captures a mindset.
Says
The Says quadrant contains (ideally) verbatim, direct quotes from user research.
Specific quotes are more useful than generic ones. A quote like “I want something reliable” could come from a user of almost any product, so it gives the team nothing to act on. A detailed quote pinpoints where the experience broke down:
“I am loyal to Delta because I never have a bad experience.”
“I don’t understand what to do from here.”
“I screenshot the confirmation page every time, because I don’t trust that the email will actually arrive.”
Thinks
The Thinks quadrant captures what goes through the user’s mind during the experience. Look at your qualitative research and ask yourself: what is the user thinking and what matters to them?
The same content can appear in both Says and Thinks. However, pay special attention to what users think, but are not willing to say out loud. Try to understand why they are reluctant to share — are they unsure, self-conscious, polite, or worried about how the answer will land?
“This is really annoying.”
“Am I dumb for not understanding this?”
Does
The Does quadrant captures the user’s behaviors — what they do as opposed to what they say. From the research, ask what the user actually does and how they go about it.
Refreshes the page several times
Shops around to compare prices
Keeps a competitor’s pricing page open in a second tab while filling out the form
If your participants use assistive technology, capture that here as well. How someone moves through a form with a screen reader or where they switch to keyboard navigation belongs in the Does quadrant.
Feels
The Feels quadrant captures the user’s emotions. Write each entry as an adjective plus a short phrase that explains it — what worries the user, what excites them, and how they feel about the experience.
Impatient: pages load too slowly
Confused: too many contradictory prices
Worried: they are doing something wrong
The context matters more than the adjective. “Anxious” on its own gives the team nothing to design against, while “Anxious: can’t tell whether the payment went through” points at a fix.
Across Quadrants
Users are complex, and it’s natural for the quadrants to occasionally disagree. You may encounter positive actions next to negative emotions.
Those contradictions are worth investigating, because they often point to something that your users did not articulate during research. It is our job as UX professionals to investigate the cause of the conflict and resolve it.
Additionally, some quadrants may feel ambiguous or overlapping — for example, it may be difficult to distinguish between Thinks and Feels. Do not focus too much on being precise: if an item could fit into multiple quadrants, pick one and move on. The quadrants exist to deepen your understanding of users and to make sure no dimension gets left out. If you cannot fill a quadrant, it’s a strong signal that you need more research before moving further in the design process.
Why Use Empathy Maps
Empathy maps can be used throughout any UX process to establish common ground among team members and to understand and prioritize user needs. In user-centered design, they are best used from the beginning of the design process.
Both the process of making an empathy map and the finished artifact have important benefits for the organization.
Capture Who Users Are
The empathy-mapping process distills and categorizes your knowledge about the user in one place. Use it to:
- Organize and make sense of qualitative data such as research notes, survey answers, and interview transcripts
- Discover gaps in your knowledge and plan the research needed to fill them — a sparse map is a sign that you need more data
- Create personas by grouping the empathy maps of similar individual users
Educate Teammates About Users
An empathy map is a quick, digestible way to illustrate user attitudes and behaviors. Once created, it should act as a shared reference point for a project and protect the team from bias and unfounded assumptions.
Keep empathy maps current by revising and adjusting them as you do more research. A map nobody has touched since the kickoff workshop will not hold up the next time someone challenges a design direction.
Collect Data Directly from Users
Empathy maps can also be used as a research instrument. At the end of an interview or usability session, hand the participant a blank template and ask them to fill in the four quadrants for the experience they just went through. Their answers give you a supplementary data source and a ready-made starting point for summarizing the session.
Also, the Thinks and Feels quadrants often surface reactions that participants did not volunteer during the session itself. Treat the Does quadrant with some caution, though: what people report doing is not always what you observed them do.
Process: How to Build an Empathy Map
1. Define Scope and Goals
Decide what user or persona you will map. Always start with a one-to-one mapping: one user or persona per empathy map. If you have several personas, build a map for each.
Define your primary purpose. If the goal is to align the team on a user, be sure everyone is present for the activity. If the goal is to analyze an interview transcript, set a clear scope and timebox the effort so you have time to map several interviews.
2. Gather Materials
Your purpose should dictate the medium. For an in-person team workshop, have a large whiteboard, sticky notes, and markers ready. If you are mapping alone, any format works as long as it’s easy to share with the rest of the team.
3. Collect Research
Gather the research that will fuel your empathy map. Empathy mapping is a qualitative method, so you need qualitative inputs: user interviews, field studies, diary studies, listening sessions, or qualitative surveys. Support tickets and call logs also work well, because they capture what users complain about unprompted.
Every sticky note on the finished map should trace back to a specific research artifact — a line in a transcript, a moment in a session recording, an observation note. If you can’t point to the source, the note is an assumption, and it should be labeled as one.
4. Individually Generate Sticky Notes for Each Quadrant
Once you have research inputs, you can proceed to mapping as a team. Everybody should read through the research individually. As each team member digests the data, they fill out sticky notes that align to the four quadrants. People should not see each other’s notes in this phase, to prevent early opinions from anchoring the group. Once the individual stickies are generated, workshop participants can then add their notes to the shared map.
5. Converge to Cluster and Synthesize
Together, the team goes through the notes on the board and clusters similar ones within each quadrant. Give each cluster a name that captures its theme, such as “validation from others” or “workarounds”; the same theme can appear in more than one quadrant. Clustering is where discussion and alignment happen: the goal is for everyone on the team to leave with the same understanding of the user.
Once your empathy map is clustered, work through it out loud as a group:
- What outliers (or data points that did not fit in any cluster) are there?
- What themes were repeated in all the quadrants?
- What themes only exist in one quadrant?
- What gaps exist in our understanding?
6. Polish and Plan
If you need more detail, adapt the map by adding quadrants — Goals is a common fifth — or by increasing the specificity of existing quadrants. Polish and digitize the output accordingly. Include the user, any outstanding questions, the date, and a version number, so a reader six months from now knows how current the map is.
Using AI in Empathy Mapping
AI tools have an obvious appeal for empathy mapping, because the method is labor-intensive and the output looks like something an AI tool could produce quickly, so the temptation to hand it over to AI is understandable. What matters is what you hand over: asking AI to organize evidence you already have is very different from asking it to generate evidence you don’t have.
Never rely on AI tools to perform all your analysis. Before using AI at any step of the process, ask:
- Can I trace this sticky note back to actual research evidence?
- If I had written this note myself, without AI, would I call it a finding or an assumption?
Common Mistakes
Mapping assumptions instead of research. A team that fills a map from memory produces a record of its own beliefs, then treats that record as evidence for the rest of the project. If you are mapping before research, label the map as a set of hypotheses and plan the study that will test them.
Treating the map as a deliverable. A map that gets exported, presented once, and filed away isn’t valuable. The value comes from returning to it when a design decision is contested and asking what the map actually supports.
Defining the segment too broadly. A map covering “our customers” ends up holding notes that contradict each other because they come from different people with different goals. If the contradictions can’t be investigated, the segment is too wide.
Letting the map go stale. Users, products, and markets change. Date the map, version it, and put a review on the calendar for the next time the team plans research.
Conclusion
Empathy maps help teams build empathy with end users. When they are based on real data and combined with other mapping methods, they can:
- Remove bias from designs and align the team on a single, shared understanding of the user
- Discover gaps in research
- Explain what drives user behaviors
- Guide the team towards meaningful innovation
Shorten Songs, Lengthen Audio, Find Loops, or Create Music from Text (Website)
Audjust AI provides a browser-based tool for trimming, looping, and generating instrumental audio tracks from text prompts.
Summary
Original Article
Warm Product Intro Bed
Warm cinematic piano, soft strings, light percussion, and no vocals. Build a clean instrumental bed for a 15-second product teaser.
Why It Works
Useful when a user wants a calm premium cue they can trim, loop, or build on for brand videos and short promos.
Creative Jobs Have Plummeted amid the AI Boom
The US creative sector has shed over 200,000 jobs since 2020, marking one of the worst employment declines in modern history.
Summary
Deep Dive
- The 200,000 job loss figure reflects a multi-year decline across media and entertainment sectors.
- The decline began around 2020 and accelerated alongside the public availability of models like Midjourney, DALL-E, and ChatGPT.
- Hiring freezes and industry consolidation have compounded the job market pressure alongside automation.
- This trend is particularly visible in copywriting, illustration, and entry-level research roles.
Original Article
Film, TV, publishing and other creative industries have shed more than 200,000 US jobs in four years, one of their worst stretches in decades.
What is IT Automation? A Complete Guide
IT automation is evolving from simple task scripting to agentic orchestration, requiring standardized governance and clear business-case metrics.
Summary
Deep Dive
- Phases: Analysis, Implementation, Integration, and Maintenance.
- Categories: Infrastructure as Code (IaC), CI/CD, Configuration Management, and AIOps.
- Challenges: Over-automation risks, lack of specialized skills, and vendor lock-in.
- Best Practices: Start with small, frequently repeated tasks; ensure stakeholder buy-in (Security, Legal, IT); and treat automation as code (version control).
Decoder
- Hyperautomation: A business-driven approach that identifies and automates as many IT and business processes as possible using AI and machine learning.
- Citizen Developer: A user who creates business applications using low-code/no-code platforms, typically without professional software engineering training.
Original Article
IT automation replaces manual data center and cloud work with repeatable scripted processes.
Amazon Introduces a Completely Redesigned Kindle Family
Amazon redesigned its entire Kindle lineup with flush-front displays, faster performance, and a new color e-reader called Colorsoft.
Summary
Original Article
Amazon has redesigned its whole Kindle family, with the Kindle, Paperwhite, and Colorsoft all thinner and lighter, using a flush-front display and edge-to-edge casing color. The 6-inch Kindle starts at $149.99 and is the most pocket-sized yet, while new accessories include a Bluetooth page-turn remote and magnetic case with page-turn buttons.
Typeface Design: PP Kyoto by Pangram Pangram
Pangram Pangram released PP Kyoto, a multilingual typeface merging Western slab-serif geometry with fluid Japanese calligraphic elements.
Summary
Decoder
- Teardrop terminals: A style of stroke ending that mimics the shape of a teardrop, often found in traditional serif typography.
Original Article
PP Kyoto is a multilingual typeface by Pangram Pangram that combines the architectural geometry of Western slab serifs with the fluid qualities of Japanese calligraphy. Heavy horizontal slabs, sweeping teardrop terminals, and prominent circular dots create tension between structural weight and organic movement, while nine weights and matching italics provide flexibility across display and editorial settings. Its 692 glyphs cover Latin, Hiragana, and Katakana, with optical adjustments and generous counterforms designed to preserve clarity across languages and sizes.
The 6 Reasons 79% of Our Design Clients Come Back
Dtail Studio sustains a 79% client retention rate by prioritizing upfront payment, unbilled communication, and honest strategic pushback over design-only deliverables.
Summary
Original Article
Good design is only the entry fee for a studio that has seen 79% of its clients since 2019 return for a second project. Retention comes from listening first, committing to a direction early, skipping per-message billing, settling payment upfront, doing unrequested work, and choosing exciting projects. Long partnerships then pay off in less re-explaining, faster and more honest feedback, a consistent direction, and room for braver choices.
Viral AI Animation Sparks Controversy over Familiar Character Design
A viral animation by Antrofrog highlights ongoing public distrust of AI-generated content due to perceived character design theft and disjointed aesthetic styles.
Summary
Original Article
A viral AI animation by artist Antrofrog drew criticism for a jumble of art styles, a strange plot, and character designs some called stolen.