Sheesh, 162 links this month. What is going ON? There is just too much interesting material being written, and that’s after filtering out clanker slop. I’m pretty sure I need to curtail the scope of these link posts at some point. I added the 🔥 emoji a while back to make it easier for folk to just skim for the top links.
It’s been a quiet blogging month for me, with just one post capturing my rant on LinkedIn about companies using LLMs to write content…to feed to LLMs. This month I’ve been mostly busy with the launch of the new blog over on Confluent Developer. The first two articles are up:
-
Keith Lee - How we cut Flink OOMKills by 91.2%: Zombie block cache, phantom CPUs (and some bonus AI Learnings).
-
Sandon Jacobs - Kafka Queues for Quarkus.
Make sure to follow the Confluent Developer Blog RSS feed for more!
|
Kafka and Event Streaming 🔗
-
🔥 Prakash Kava (Atlassian) - Transitioning from Kinesis to Kafka for 145 Billion Daily Events.
-
Josh Parsons - Transforming How We Run Kafka at Honeycomb.
-
Yaroslav Tkachenko discusses Apache Iggy, S2, and OpenData Log: The New Wave of the Streaming Log Technologies.
-
Good performance analysis from Jack Vanlightly (as always!) looking at Kafka’s
linger.mssetting. -
Mateusz Gołąbek (SoftwareMill) - Eligible Leader Replicas (ELR) in Kafka 4.1.
-
Julien Brunet (Michelin) - Kafka Consumption, Made Simple and Safe.
-
Aiven’s Olena Babenko argues that Apache Kafka Share Groups are NOT true queues (and posits that it’s a good thing).
-
🔥 I’m thoroughly enjoying the wave of excellent content from Grzegorz Kocur and Michał Matłoka on their Monedula site. This month they have a look at KIP-714 at Kafka client metrics over OTLP, as well as an update to their Kafka Simulator v1.1: Understanding Kafka Producer Semantics.
-
Richard Artoul at WarpStream looks at performance improvements on GCP, with One Weird Trick to Make Rapid Storage 40x Faster.
-
Varun Gandhi argues that Job queues are deceptively tricky.
-
Jonas Geiregat - Putting a Kafka Topic Naming Convention into Practice with Terraform.
Stream Processing 🔗
-
🔥 Valentin Touffet and Alexandre Olivier describe in great detail How Datadog measure data completeness at scale.
-
Nicoleta Lazar (Fresha) - Processing Patterns with Apache Fluss.
-
Apache Fluss has graduated from the incubator to become a Top Level Project (TLP) at the Apache Software Foundation (ASF).
-
🔥 Aleksandr Birin (Zalando) - From Homegrown to Flink: Migrating a Stateful Ad Event Join at Scale.
-
Katya Gorshkova - Hands-On with Flink — Part 7: Exploring Flink Checkpoints. (Previously parts: 6, 5, 4, 3, 2, 1)
-
Sébastien Viale and Loïc Greffier (Michelin) - Dead Letter Queue in Kafka Streams (KIP-1034).
-
🎥 Recording of the LinkedIn Stream Processing Meetup - June 24, 2026, with three talks:
-
Forkable Shared Logs: Enabling Testing, Analysis, and Agentic Workloads on Real-Time Data — Shreesha G. Bhat, UIUC.
-
Modernizing Flink Jobs at Scale: A Platform Approach to Flink 2 Upgrade — Daniel Trager, Mark Cho, Netflix.
-
Flink Issues Classification Engine — Manan Chandra, Stuart Tsao, Ankitha Gavinolla, Michael Barskii, LinkedIn.
-
-
Lalith Suresh (Feldera) was a guest on Kris Jenkins' Developer Voices podcast, explaining Incremental View Maintenance (IVM): What If Every SQL Query Could Update Incrementally?.
-
A profusion of new projects this month, including:
-
StreamFusion - an open-source Flink Accelerator built on Apache DataFusion.
-
flare-db - an Apache Beam native streaming database built in Rust.
-
cobble - an embedded key-value store designed for use in distributed systems and standalone applications, including as a state backend for Flink.
-
BlazeRules - Vectorized YAML-Decision Engine.
-
Str:::lab Studio - A zero-dependency, modular browser SQL Studio for Flink SQL.
-
Analytics 🔗
-
Junhyun Ko and Byeongwoo Lee (Rapport Labs) - How we adopted StarRocks for ad performance data.
-
Rubens Minoru Andako Bueno takes a look at Apache Doris for real-time analytics.
DuckDB 🔗
-
🔥 Kyle Cheung with part 2 of their series on DuckDB Internals and what makes it so darn fast, this time looking at how it uses Vectorized Execution. Find part 1 here.
-
DuckDB is growing in adoption, not just on developers' laptops but as the platform on which companies build analytics:
-
Arcesium moved from Trino (and previously to that, Athena) to DuckDB, storing their data in Iceberg on S3.
-
PostHog moved from ClickHouse to DuckDB for their data warehouse.
-
-
🔥 David Gasquez has published a useful list of Awesome DuckDB resources.
-
Massimo Meneghello - the-stats-duck v0.6.0 — statistics that live in your SQL.
ClickHouse 🔗
-
ClickHouse can now run in your browser with chDB SQL Shell (kinda like DuckDB Wasm has been able to for a while 😉).
-
Henry Haefliger (Momentic) describes why they moved from Postgres to ClickHouse.
-
Zepto Tech - How We Built Ads Analytics That Actually Works using ClickHouse.
-
Why? Because You Can!
Data Platforms & Architectures 🔗
-
Reynold Xin (Databricks) - From monolith to Lakebase to LTAP: rethinking the database from storage up.
-
Almost as certain as a vendor publishing benchmarks favourable to themselves is another vendor writing a take-down of said benchmark. Melvyn Peignon at ClickHouse does just this, arguing that the results that showed ClickHouse in an unfavourable light compared to Databricks were not reproducible and thus unfair. I also learnt from this post my new favourite word: Obscurantism.
-
More fallout from Databricks' Reyden benchmark, this time from Rudi Leibbrandt at Snowflake, arguing that real workload performance is the metric that matters.
-
Will Edwards (Spotify) describes how they use "Random Access Parquet" (RAP) for accelerating queries on files held in their data lake.
-
Sarp Kaya - Your gRPC Service as a SQL Table: Building an Interoperable Trino Connector.
-
Chris Gambill - The Modern Data Stack Hoax.
-
A couple of useful technology and tooling roundups, always handy for folk newer to the scene:
-
Gleb Mezhanskiy - The Modern Data Stack: Open-source edition.
-
Oleh Korniienko - Guide to data tools landscape for developers.
-
-
Chad Sanderson - The Shift Left Manifesto - v2.
-
Ananth Packkildurai - Anatomy of a Privacy-Safe Data Platform.
Data Engineering and Pipelines 🔗
-
Kirill Bobrov - Most of Your Backfills Didn’t Have to Happen.
-
René Luijk - dbt Incremental Models vs. Snowflake Dynamic Tables.
-
Amit Prabhu - How We Refresh Razorpay’s Data Warehouse 10x Faster with Graphs and Indexes.
-
Our colleagues in software engineering have long had LeetCode - now we have LeetData for SQL, PySpark, and Data Engineering interview practice.
CDC 🔗
-
🔥 Chris Cranford - Oracle Log Mining Simplified in Debezium 3.6.
-
Mario Fiore Vitale - Building Kafka-Less Data Integration Pipelines with Debezium.
-
Philippe Camus - Running the Debezium Platform on AWS: PostgreSQL to Amazon Kinesis.
-
Jiufeng Liu (Intuit Credit Karma) - Streaming CDC at Scale.
Data Modelling 🔗
-
🔥 Joe Reis - The Database Is Not the Data Model.
-
Connor Charles (AutoTrader) - Unifying SQL Transformation Logic with a Custom Semantic Model.
-
Apache Ossie has been launched as an incubating Apache project, coming from the Open Semantic Interchange. It defines itself as:
an open specification that defines a vendor-neutral format for expressing business metrics, dimensions, and their relationships. It enables any tool to consume and produce semantic definitions without loss of meaning.
Open Table Formats (OTF), Catalogs, Lakehouses etc. 🔗
-
🔥 Gunnar Morling - A Fast Path for Fixed-Length Lists in Parquet.
-
Angel Conde - Debugging a Native Memory Leak in Apache Iceberg v3.
-
Chad Lagore (Affirm) writes about Time Travel and Data Validation in their new platform built around Aurora MySQL, CDC, and Iceberg.
-
Artur Borycki (Teradata) - Puffin-Backed Vector Indexes: Attaching Approximate Nearest Neighbor Indexes to Apache Iceberg Snapshots for Compute-Disaggregated Query Engines.
-
🔥 Francisco Morillo has published a reproducible, open-source benchmark of Apache Spark 4.1 and Apache Flink 2.2 streaming 100,000 records/second into Apache Iceberg.
-
Rahul Penti (Grab) writes about the adoption of Apache Iceberg for their data lake, moving away from Hive and Parquet.
-
Two interesting posts from folk at Snowflake. More interesting for me than the actual conclusions (which, given their employer, will either be condemned or condoned by the reader’s place in the dbx/sf ecosystem) is the illustration of the sheer complexity still involved in plumbing some of these things together, providing a useful roadmap for people as and when they are making these evaluations and decisions.
-
Paul Needleman looks at the interoperability between Databricks and Snowflake and in particular the role that Catalogs play here.
-
Bharath Suresh evaluates what does and doesn’t work when using Iceberg from Snowflake.
-
-
Ravi Dubey did some interesting experiments with Amazon S3 Tables and where they do—and don’t—work well.
-
Kushal Kumar Vishwakarma - When Should You Use Liquid Clustering in Delta Lake?.
-
Apache Gravitino 1.3.0 has been released.
-
Jack Vanlightly - Benchmarking Hardwood 1.0 on a Threadripper 9980X.
-
🔥 Inspect Parquet files in your browser using hardwood dive, built using Hardwood and WebAssembly (Wasm) with GraalVM Web Image.
RDBMS 🔗
-
Aurora DSQL: Scalable, Multi-Region OLTP is a new paper from Marc Brooker et al. There’s commentary from Marc himself, as well as Murat Demirbas.
-
Rahul Bansal and colleagues at Affirm share some lessons for scaling Aurora MySQL for peak traffic.
-
🔥 File under "literally impossible until LLMs came along, and now relatively trivial": someone has rewritten Postgres in Rust. Michael Malis has published pgrust, which targets compatibility with Postgres 18.3 and matches Postgres’s expected output across more than 46,000 regression queries. There are some detailed posts going into how they did it here:
-
Also LLM-created is Nikolay Samokhvalov’s PGSimCity. I can’t decide if I love this, or not. It’s a very clever idea, and looks great. To me it does also illustrate the problem of an LLM building literally what it’s told. The UX is arguably too busy, too dense, too much to comprehend—which is ironic given that the purpose of what has been built is to educate and clarify.
-
Burak Sen - Reading The Internals of PostgreSQL: Database Cluster, Databases, and Tables.
-
Haki Benita - How to Achieve Pruning When Querying by Non-Partitioned Columns in PostgreSQL.
-
Alexander Belanger - A guide to preventing Postgres from toppling over.
-
Andrei Ivaniuk - PostgreSQL on AWS: Size & Benchmark EC2 Instances.
-
🔥 Mark Lin and Patrick Cosmo (Warner Music Group) - Rethinking NoSQL: Why We Migrated From Cassandra to PostgreSQL RDS.
-
Chris Pederkoff (Airtable) - Instant Schema Changes in Airtable’s New Database.
General Data Stuff 🔗
-
Qi Zhu and Andrew Lamb - Optimizing for Almost Sorted Data: Sort Pushdown in Apache DataFusion.
-
🔥 Sem Sinchenko - Algorithms on billion-scale graph using 10GB RAM: I love DataFusion!
-
Mark Callaghan - Orasort: Common prefix skipping, adaptive sort.
-
Kyle Davis - The secret life of data in Valkey.
-
Jon Avezbaki - notes on the Designing Data-Intensive Applications book.
-
Oskar Dudycz - Addition by subtraction in software design.
Career Advice 🔗
Not a regular section, but I read a bunch of posts this month that I would love to have read when I was earlier in my career.
-
Chris Hillman - You Can’t Incentivise a Pipeline That Doesn’t Break.
-
Swizec Teller - How to be useful as a software architect.
-
🔥 Sean Goedecke - What does "playing politics" mean for software engineers?
-
🔥 Dan Moore - Ask for no, don’t ask for yes.
-
Sahar Massachi, Zach Wilson - Junior data engineers build pipelines. Seniors build trust.
-
Rashi Desai - Lessons learnt from working in consulting, and staying relevant in your career in the AI age.
AI 🔗
I warned you previously…this AI stuff is here to stay, and it’d be short-sighted to think otherwise. As I read and learn more about it, I’m going to share interesting links (the clue is in the blog post title) that I find—whilst trying to avoid the breathless hype and slop.
-
If you only read one piece this month, read this one: 🔥 Jay Acunzo - The best response to AI slop, infinite advice, and online noise is from Robin Williams.
Big picture, hype, and economics 🔗
-
🔥 I’m a huge fan of Benedict Evans and his calm, measured, and informed commentary on WTaF is happening in the industry at the moment. Check out his rational conversation on where AI is actually going with Lenny Rachitsky, as well as articles on Ways to think about token pricing, and Predicting AI job exposure.
-
Cory Doctorow - AI solipsists and AI cynics.
-
🔥 Nikhil Suresh - AI Mania Is Eviscerating Global Decisionmaking.
-
rruxandra - Are the LLM Wars the Database Wars?
Careers, culture, and slop 🔗
-
Inspired by WarGames, I wrote a short piece: THE ONLY WINNING MOVE IS NOT TO PLAY.
-
Noam Segal, Lenny Rachitsky - How tech workers are feeling in 2026: a workforce splitting in two.
-
Addy Osmani - The Agent-Era Career.
-
Kirill Bobrov - My Empathy Is a Sticker Detector.
-
Having tried to navigate the issue of AI-assisted content and straight-up AI slop as mod on r/apachekafka, I was interested to see how the HN crowd would deal with it: Ask HN: Add flag for AI-generated articles.
-
Chris Ford and Richard Gall (Thoughtworks) - The zero-cost fallacy: Open source software in the agentic era.
-
Roland Huß - Open Source as we know it is dying a slow death (LinkedIn post, and thread).
-
Jamie Dobson - Is Open Source Dead?
-
Yes, yes: Mr Betteridge would like a word…
-
The OpenAI / Hugging Face incident 🔗
This one is straight up bonkers. tl;dr: A new model that OpenAI were testing (supposedly in a 'sandbox') decided the best way to solve the test it’d been given was to break into Hugging Face, which it duly did. These articles go into a lot more of the detail, analysis, and commentary.
-
🔥 Casey Newton - A big week for AI denialism.
-
The two disclosures: OpenAI’s and Hugging Face’s.
-
Simon Willison - OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened.
-
Zvi Mowshowitz - OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation.
-
antirez - The real AI risk is inside the labs.
AI in Data 🔗
How we’re architecting and building our data systems is changing. Some of it’s hype…some of it’s real.
-
🔥 Joe Reis - To Every Agent Its Own Database.
-
Sven Balnojan - Agentic analytics is bullshit. It saves your ass.
-
🔥 Gwen Shapira - Postgres for Production Agents: Your Relational Foundation for Enterprise AI.
-
Rituparna Das - Three Ways I Actually Use AI in Data Viz.
-
Aditya G. Parameswaran and a team from Berkeley including Joseph Hellerstein and Ion Stoica have published an article: Intelligence is Free, Now What? Data Systems for, of, and by Agents.
-
PostHog continue to make huge moves away from simply providing analytics and to going all-in on both the benefits that AI can provide engineers, and building a data platform to enable that. They moot the idea of the context warehouse (as an evolution of the data warehouse), although I can’t quite decide if the article is positioning a new architectural concept or simply a product that they’re offering (maybe both?).
-
Mahesh Kshirsagar (Meta) - How We Built DEmate: Taming LLMs for Data Engineering.
-
Details of how Expedia use LLMs to analyse Spark SQL plans.
-
Vikas Rai (Halodoc) - How We Turned Data Engineering Runbooks Into Reliable AI Skills.
AI in Software Engineering 🔗
-
🔥 Vinay Gaba and team at Databricks - Benchmarking Coding Agents on Databricks' Multi-Million Line Codebase.
-
Liz Fong-Jones - 30 to 70 PRs a Day: How We Managed to Not Wreck Our Systems.
-
Gergely Orosz - What is "loop engineering?"
-
Gregor Ojstersek - How to Do Spec-Driven Development.
-
🔥 Jamie Brandon - Artificial adventures.
-
Birgitta Böckeler has a couple of good posts about using local models:
-
Adam Friedmann (Wix) ran a bunch of tests to see how useful and important Skills are in agentic coding (it turns out, not always).
LLMs in products and platforms 🔗
-
Shuai Guo - Stop Choosing Between Local and Cloud LLMs: A Field Guide to Hybrid Patterns.
-
ByteByteGo - Best Practices for Building AI Agents That Work in Production.
-
🔥 Paul Iusztin - Agent Memory From Scratch.
-
Kendrick Tan and team - How we help Grab build and run AI agents at scale.
-
Fabio Flores, David Lindsay, Steven Xu, and colleagues at DoorDash with a series of deep-dives about their shared agent platform, evals, and using LLM juries, multi-modal AI, and more for building food metadata.
-
Doron Shapira (Wix) on building a dashboarding tool with AI that then gets adopted by the whole company.
-
Baharak Saberidokht (Airbnb) - From weeks to a day: how we made LLM evaluation fast enough to iterate on.
And finally… 🔗
Nothing to do with data, but stuff that I’ve found interesting or has made me smile.
Lead 🔗
-
Couple of great articles from Michael Lopp (a.k.a. Rands) - So You Want to Fix Your All Hands and Entropy Crushers.
Blog 🔗
-
🔥 Jim Nielsen - Blogging Can Just Be Stating The Obvious.
-
Julia Evans - Write for 1 person.
-
Michael Lopp - The Motivation.
Nerd 🔗
-
I’m just fascinated by the idea that I can run a ZX Spectrum 48K in my web browser. That, and dozens more, are here: Tiny Emulators.
-
More nostalgia at the Winamp Skin Museum.
-
Randy Au - So how does lightning locating work?
-
🔥 Very cool security deep-dive from Appaji Chintimi digging into fake job interview git hook malware.
Listen 🔗
-
🔥 Dan Carlin’s Hardcore History podcast is like no other. I’ve recently had the absolute pleasure of listening to the first three parts of his Mania for Subjugation series about Alexander the Great. Check it out—they are just fantastic:
Quote 🔗
Not technically links…but my blog, my arbitrary rules :P
Have you ever noticed that anybody driving slower than you is an idiot, and anyone going faster than you is a maniac?
You cannot reason a person out of a position he did not reason himself into in the first place.
|