How we knew COVID was over (and what our models had to unlearn) (11 minute read)
Model maintenance is not just retraining. When COVID-era demand patterns faded, Airbnb's forecasting team separated parameter drift from structural change, choosing among refit, respecify, or hold decisions instead of blindly updating thousands of markets. The result was a useful operating model for production forecasting under regime shifts.
|
From 17ms to 0.04ms: How to Design the Right SQL Index (6 minute read)
Good SQL indexes come from real query patterns, not the table schema. Composite indexes need the right column order, usually equality filters first and sort or range columns after, with EXPLAIN ANALYZE used to verify the result. The right indexes cut example queries from 17ms to 0.04ms and 436ms to 0.5ms.
|
A tale of two Flink autoscalers (7 minute read)
Netflix is moving 30,000+ Flink jobs from a homegrown autoscaler toward the Apache Flink Autoscaler. The old system saved 25–45% by scaling whole clusters from container metrics, but missed operator-level state and suffered from telemetry gaps. The OSS autoscaler reasons from inside each job and produced 58% annualized compute savings for one team.
|
Zero-sum by design: 10 years of Uber's payments platform (8 minute read)
Uber's Gulfstream payments platform now supports $217B in annualized gross bookings, nearly twice that in money movement, and ledger balances for 1.2B+ entities. The design relies on immutable zero-sum money orders, double-entry accounting, DynamoDB-backed balances, Kafka pipelines, Cadence for synchronous flows, and generic payment-instrument interfaces that let new business lines plug in with minimal core changes.
|
|
The most valuable API in the modern data stack returns no data (10 minute read)
Bulk export APIs are giving way to shared table references. Iceberg REST catalogs, object storage, Parquet, and credential vending let producers publish metadata while consumers bring their own compute, avoiding pagination, rate limits, and duplicate pipelines. The hard parts move to contracts, compaction, semantics, privacy, and table maintenance promises.
|
Not every problem needs an AI agent (5 minute read)
Production AI systems work best when each component earns its complexity. Recommenders, moderation, warehouse queries, and search cold starts need different mixes of deterministic logic, classic ML, semantic layers, and LLMs. The strongest pattern is often hybrid: use agents for reasoning, but keep business logic, metrics, routing, and guardrails deterministic where possible.
|
|
DuckDB v2.0: Your database deserves a better parser (9 minute read)
DuckDB v2.0 replaces its PostgreSQL-derived SQL parser with a PEG-based parser while preserving the downstream binder and AST pipeline. The change removes Bison conflict pain, avoids exponential backtracking on malformed input, and enables runtime grammar extensions that can add syntax without reimplementing SQL parsing.
|
TrueFoundry open-sources TrueForge, an enterprise AI agent harness (8 minute read)
TrueForge is a self-hostable agent harness focused on reducing task cost through context compaction, delayed tool-schema loading, subagents, and sandbox-as-a-tool execution. The release matters for data teams experimenting with governed agents because it separates orchestration from model choice and pairs naturally with gateways for identity, access control, observability, and budgets.
|
|
Using Models to Create Models of New York City (20 minute read)
Peter Sobot used 1,000+ photos, Claude, and Gaussian Splatting to build a 3D model of the New York skyline. Running 74+ subagents and roughly 280 experiments, the project showed that AI can handle much of the technical workflow, but still struggles with visual judgement, messy data, and knowing when the result is genuinely good.
|
|
|
|
|
Comments
Post a Comment
VHAVENDA IT SOLUTIONS AND SERVICES WOULD LIKE TO HEAR FROM YOUπ«΅πΌπ«΅πΌπ«΅πΌπ«΅πΌ