In the UK August is known in the press as "silly season" (because nothing of note happens, so the only stories published are nonsense ones). Turns out this doesn’t apply to the world of data and AI blog posts - still over 50 interesting ones for me to share with you this month.
It’s been fun to see the interest in these blog posts increasing - I took a moment to check the stats and there’s a definite up-and-to-the-right pattern which I like :)
The subscriber base for the x-posts of this that I do over at Substack is growing in a healthy manner too since I started it last December:
|
Kafka and Event Streaming 🔗
-
Michał Matłoka - Kafka Simulator v1.3 — durability and replication.
-
Patrick Hamann and Mike Fisher describe their successful migration from Google Pub/Sub to NATS.
-
Grzegorz Kocur - Managing Kafka with AI and MCP: A Hands-On mcp-confluent Guide.
-
Atul Adya and colleagues at Databricks published a paper last year looking at the limitations of pubsub systems.
-
🔥 Apache Fluss is now a Top-Level Project at the ASF, and the project has published a case study on their blog about how RedNote moved from Kafka to Fluss for real-time indexing.
-
Editor’s note: Fluss is a funny one for my existing set of categories. It’s not
Event streaming, it’s notstream processing; it’s storage & serving, I guess. I need to think this one through as the project continues to gain adoption, perhaps :)
-
Stream Processing 🔗
-
More good content in Katya Gorshkova’s "Hands on with Flink" series: Part 8: Flink SQL in Application Mode. (Previous parts: 7, 6, 5, 4, 3, 2, 1).
-
🔥 Yogesh Nagarur shares details of how Netflix adopted streaming joins in Apache Flink: Evolving Netflix’s Ads Event Pipeline for Live — Part II (you can find part I here).
-
Mohamed Moataz El Zein - Scaling a stateful exactly-once Flink job to 300M+ RPM.
-
Dale Lane - Running Flink jobs in MiniCluster using the Kubernetes Operator.
Analytics 🔗
-
🔥 Exciting - DuckDB 2.0 is on its way! Mark Raasveldt and Hannes Mühleisen share a preview.
-
Jovan Stojiljkovic has a nice analysis comparing SQLite vs DuckDB (spoiler: DuckDB FTW).
-
Matt Martin compares Spark vs DuckDB for the same task: Querying 1 Thousand JSON Files From S3.
Data Platforms & Architectures 🔗
-
Nimish Sheth and colleagues at Uber look at 10 Years of Uber’s Payments Platform, whilst elsewhere in the company Daniel Musgrave and team describe Simplifying Data and Product Integrations with a Data Abstraction Layer.
-
🔥 Fascinating explanation from Nilesh Mishra and Ajit Koti of the kind of graph queries they serve—and how—in part 3 of the series about Netflix’s Real-Time Distributed Graph (part 1, part 2).
-
Oleksii Tkachuk and team at Netflix describe how they optimised handling the cold tier of temporal event data at scale. There’s also a related article from 2024 about the same platform.
Data Engineering and Pipelines 🔗
-
🔥 Richard Wilmer - Real-time Personalisation on Autotrader.
-
Thijs Nieuwdorp - Prototype on a laptop, scale to 16 billion rows: one Polars query.
-
Thiago Rocha Salvatore and Lizzie Epton at PostHog describe their approach to the semantic layer.
-
Brad Coles - Data Engineering’s Shift from Imperative to Declarative.
CDC 🔗
-
Robert Oliveira documents some design decisions taken in their move from batch snapshots to near-real-time data.
-
A hands-on guide from the StarRocks Engineering team on how to do CDC from Postgres to StarRocks.
-
Mukesh Agrawal and colleagues at AWS show how to do CDC from Aurora DSQL into Apache Iceberg.
-
Andreas Andreakis argues that Change-Data-Capture Doesn’t Solve Dual-Writes, supported by a paper that he’s published: Machine-Checked Dual-Write Recovery from a Committed Log.
Open Table Formats (OTF), Catalogs, Lakehouses etc. 🔗
-
🔥 Gunnar Morling - A Fast Path for Fixed-Length Lists in Parquet.
-
A couple of Apache Iceberg posts from the Guidewire Engineering Team, covering table maintenance and concurrency challenges.
-
A neat little utility from the folk on the Nessie project: iceberg-catalog-migrator can be used to bulk migrate tables between catalogs without a data copy.
-
Details from Pankaj Mohapatra and colleagues at Uber about how they use column stats in Apache Hudi to optimise scans over large volumes of data: Running Cost-Efficient Export Workloads at Uber.
-
Neelesh Salian writes about the support for the Variant type in Iceberg.
-
Yaniv Zalach has written IceGraph, a rather neat interactive Apache Iceberg debugger. You can see a live demo here.
RDBMS 🔗
-
🔥 Alex Chan (Tailscale) - How we tracked down a 16-year-old SQLite bug.
-
It used to be simple. A database over here for your transactional workload, one over there for all your heavy analytics, and perhaps a few others scattered around for the cool kids playing with NoSQL or graphs etc. Over the past few years several databases have started to stretch their support to encompass other workloads. Rick Houlihan from Oracle considers a possible definition for a converged database.
-
sqlfmt is a SQL gofmt-style formatter.
-
Emilie Noel at Shopify writes about how they moved from Redis to MySQL to support the inventory reservation process at checkout under high volume.
General Tech Stuff 🔗
-
🔥 A cracking post from Joe Reis: The 'X is Dead' Fallacy.
-
Melkey Moksyakov, Janos Szathmary, and Andrew Healey share details of how Vercel migrated from Redis to DynamoDB.
-
Pragya Mehta and Sai Samant at Stripe write about how they use graph search and state machines to auto-remediate their 2000-shard MongoDB deployment.
-
A while ago this would have seemed like science fiction, but Isaac Wong at Cockroach Labs has a serious article considering What Orbital Computing Means for Distributed Databases. Whilst part of it is obviously just highlighting how awesome CockroachDB would be in many scenarios, I found it interesting nonetheless.
-
Farid Zakaria - The Mean Means Nothing.
-
Dan McKinley - Choose Boring Technology.
-
Gergely Orosz, Ivan Klaric, and Jesse Spevack have a good deep-dive article on Software engineering at a proprietary trading company: Optiver.
AI 🔗
|
I’m retiring the above disclaimer, which has been present since last autumn. If you’re still sceptical about AI and don’t believe it’s going to have a fundamental impact on how we do stuff, you’re possibly not my intended audience for these blog posts. |
AI in Data 🔗
How we’re architecting and building our data systems is changing. Some of it’s hype…some of it’s real.
-
Simon Green - Your Next 'Access Database Problem' Is an AI Agent.
-
InfoQ’s Leela Kumili has a nice summary of how Grab are using AI agents to automate analytics workflows.
-
Isobel Scott (Etsy) - Kafka App? There’s a Skill for That.
AI in Software Engineering 🔗
-
🔥 Markus Eisele - Spec-Driven Development Needs an Exit Strategy.
-
🔥 Patrick Debois & Baruch Sadogursky - The Right 300 Tokens Beat 100k Noisy Ones: The Architecture of Context Engineering.
-
Dachary Carey - A Skill is More than Markdown.
-
Ian Vanagas - This post will save you tokens.
And finally… 🔗
Nothing to do with data, but stuff that I’ve found interesting or has made me smile.
Think 🔗
-
🔥 Brad Stulberg - A "Balanced" Life is Boring.
Fun 🔗
-
Antoine Mayerowitz - Mario meets Pareto.
-
Ryan Farley - Gateway 2000’s descent from awesome to bad ads in the 90s, Part I & Part II.
Work 🔗
-
🔥 Joe Martin - The most common questions about developer marketing, answered.
-
Niklas Gruhn - Don’t be a meat proxy.
|