Sparking New Ideas for the Data Platform

3 minute read

Published:

I have been sitting with an idea for a while now, and I think it is time to write it down instead of just turning it over in my head.

The Cost Problem Every Big Data Platform Shares

Almost every enterprise company I have talked to or worked with runs into the same wall once their data platform matures: cost. It does not matter if it is Databricks, Confluent, DataIKU, or Azure Synapse β€” these platforms are powerful, but that power comes bundled with clusters, compute units, storage tiers, and licensing that scale in a direction that is hard to predict and even harder to control. Teams end up paying for capability they only use a fraction of, because the platform is built to handle everything at once rather than the specific thing your organization actually needs today.

That is not a knock on these platforms; they solve real problems at real scale. But β€œbig platform for everything” is a very expensive default when most organizations do not need everything, they need a handful of things done really well.

The AI Era Removes an Old Constraint

For years, the reason companies bought into these all-in-one platforms was simple: they did not have the developer capacity to build and maintain something more tailored. Custom tooling was a luxury reserved for organizations with large engineering teams.

I think that constraint is no longer a hard blocker. In the AI era, a small team β€” or even a single developer β€” can move faster and cover more ground than an entire team could a few years ago. The organization no longer needs a large headcount to β€œdo big things.” It needs the right small things, built well, and AI closes most of the gap that used to require that headcount.

The Idea: Small Pieces, Big Outcomes

This is where my idea starts. Instead of adopting one massive platform meant to do everything, what if the data platform was made of small, purpose-built pieces β€” each one lightweight, focused, and cheap to run β€” that together deliver the same outcomes big platforms promise, but without the overhead?

The core of it is rethinking how data and files are structured in the first place. Big platforms tend to assume you need their engine to make sense of your data. I think a more efficient structuring of data and files can remove a lot of that dependency, letting small components do the heavy lifting that today gets outsourced to a big, expensive cluster.

It is a completely different mental model from the big data platforms we are used to today. Instead of β€œone platform to rule the pipeline,” it is β€œmany small, efficient pieces, composed to deliver the big result.”

Where This Goes From Here

This is still early β€” more of a spark than a blueprint. But it is the same instinct that got RepoDB started: a thin, efficient, transparent layer can often do more than a heavy one, at a fraction of the cost. I want to explore whether that same philosophy holds true at the data platform level, not just at the data access level.

If this resonates with you, or if you have run into the same cost wall with your own data platform, I would love to hear about it. This is the kind of idea that gets sharper the more it gets challenged.

This post contains what is on my head, but it is also refined by AI.