Does Your Company Need a Data Platform? Why and When
What is an enterprise data platform?
When people talk with me about data platforms, the most common question is actually not “how do we build one?” It’s “do we really need one?” To be honest, the answer depends a lot on the situation of the company.
First, let me explain what I mean by this term. A data platform is not a new idea. Sensors on a factory line generate data all the time, and a sales company is generating sales data every second of the day. The need to collect and process data has existed for a long time. When I say “enterprise data platform”, I mean the shared foundation that collects data from all over the company, cleans it up, stores it in one place, and hands it to whoever needs it, whether that’s a finance analyst, a machine learning model or a customer-facing app. It’s not one database or one dashboard tool. It’s the whole chain: getting data in, transforming it, storing it, controlling who can see it, and serving it to lots of different users at once.
Building it is a big investment, both in money and in people. So in this post, I will put the “how” aside and focus on the two questions that should come first: why does a company need one, and when is the right time to build it?
Why companies need a data platform
What the business actually needs from data
Every time I join a new company, I like to ask people what they want from data. Interestingly, the answers almost always fall into the same four categories, no matter how big the company is or which industry it is in.
Making better decisions. This is classic business intelligence (BI): standard reports, dashboards, KPIs, and analysis for finance, sales, HR and marketing. The leadership team wants to see how the business is really doing, and make decisions based on numbers instead of feelings. A big part of this is having one version of the truth. I have joined many meetings where two teams brought two different numbers for the same metric, and in the end the whole meeting became a discussion about whose spreadsheet was correct. This happens in companies of every size.
Running the business more smoothly. A lot of data work is about day-to-day operations, not big strategic decisions. Think inventory levels, delivery times, support ticket backlogs, or the weekly report that someone spends every Monday morning putting together by hand. In fact, every company I have worked for had some kind of “Monday report”. A platform can automate this kind of work, and also send clean data back into the tools people already use, like the CRM or the marketing system, so they don’t need to search for it.
Building better products and customer experiences. This is where data mining and machine learning come in: personalized recommendations, demand forecasting, predictive maintenance, supply chain optimization, and so on. Some companies go one step further and turn data into a product itself, like analytics features inside their app, or data extracts they deliver to clients. In all these cases, the result is not just a report. The insights go back into the business automatically.
Staying compliant and managing risk. This one is less exciting. But from my experience, it is often the reason the budget gets approved, especially in bigger companies. Depending on your industry and where you operate, regulations like GDPR, CCPA, HIPAA or SOX mean you need to know what data you have, where it lives, who can see it, and how long you keep it. Auditors also want numbers that you can prove. Besides, fraud detection and internal risk control are basically data problems too.
Why not just give each team its own tools?
You could meet each of those needs separately: a BI tool for finance, a script for the ops report, a separate pipeline for the data science team. Most small companies start this way, and honestly, it works for some time. I did the same thing before. But when more and more teams need data, these separate solutions start to cause problems for each other. A shared platform can give you some things that separate tools cannot:
- One source of truth. Metrics are defined once, so “revenue” or “active customer” means the same thing on every dashboard.
- Connected data. The interesting questions usually cross team lines, like how marketing spend relates to support tickets and churn. You can only answer them if the data lives together.
- Less duplicate work. Data gets ingested and cleaned once, not five times by five teams in five slightly different ways.
- Self-serve access. People can find and query data themselves instead of filing a ticket and waiting for an engineer.
- Governance in one place. Access control, privacy rules, retention and audit trails are set up once and applied everywhere, instead of being spread across many tools that nobody fully controls.
The hidden cost of not having one
The cost of a data platform is easy to see: there is a bill from the vendor, and there is a team. The cost of not having one is hidden in many places and easy to miss. But from what I have seen, it is often bigger:
- Analysts often spend more time finding and cleaning data than analyzing it
- Engineers get pulled into one-off data requests instead of building the product
- Meetings get spent arguing about whose numbers are right
- Decisions get made on stale or incomplete data, or on nothing at all
- Every new regulation or audit becomes an emergency
- The ML project that looked great in a demo never makes it to production, because there’s no reliable data feeding it
None of these costs appear in the budget, and that is exactly why people easily ignore them.
Also, this is not only a problem for big companies. Actually, some of the most serious data issues I have dealt with were in smaller companies. When there is no platform and nobody owns the data, the cause is usually something very simple: a script that quietly stopped running, or a spreadsheet formula someone changed. Wrong numbers can stay in reports for weeks before anyone notices. And because the raw data was often not saved anywhere, there may be no way to go back and fix the history.
Why this is a bigger deal now than it used to be
Companies have been building data warehouses for more than 30 years, so what is different now? In short, there is much more data, it comes from many more places, and many more people want to use it. Even a small company now often uses dozens of SaaS tools, each holding a slice of its data. With machine learning and AI, data is no longer only something for reports. It becomes something your product depends on.
At the same time, building a platform has become much cheaper. A data warehouse used to mean buying expensive servers and software licenses, sometimes even a special hardware appliance, and hiring a team to run it. For most smaller companies, this was too much. Today, cloud warehouses, lakehouses and managed ingestion tools let a small team get a solid platform running in weeks, and many of these services charge by how much you use. The need goes up, and the cost goes down. That is why the data platform changed from something only big enterprises had, to something most growing companies will need sooner or later.
When does a company need a data platform?
Signs you’ve outgrown your current setup
When a company is small, low-code tools can cover most of the data needs. Honestly, Excel can already do a lot. In my experience, no company builds a data platform just because it sounds cool. They are usually pushed by problems that become worse and worse. Below are the signs I usually look for:
- Different teams report different numbers for the same metric, and nobody can explain why
- Analysts spend more time hunting for and cleaning data than actually analyzing it
- Simple questions like “how many active customers did we have last quarter?” take days to answer, because someone has to pull data from three systems by hand
- Reports run directly against production databases and slow down the actual product
- Important data is stuck in log files, spreadsheets on someone’s laptop, or a vendor system nobody can query
- Every team builds its own little pipeline, so the same data gets copied and transformed five different ways
- When an auditor or customer asks where a number came from, nobody can tell them
If two or three of these sound familiar to you, it is worth starting the discussion. If most of them sound familiar, you are probably already late.
Events that usually trigger it
Sometimes the push does not come from slowly growing problems, but from a specific event. Below are the ones I have seen many times:
- Rapid growth. More customers, more products and more teams means more data and more people asking questions. Something that works for 50 people often stops working for 500 people. Usually, the first thing that breaks is the one person who knows where all the data is.
- A merger or acquisition. Suddenly there are two of everything, two CRMs and two finance systems, and leadership wants one combined view of the business.
- A new regulation or audit. Requirements like GDPR or SOX force you to know exactly where your data is and who can access it.
- Launching an ML or data product. Models and customer-facing data features need reliable, fresh data, not a monthly export.
- Moving to the cloud. If you’re migrating systems anyway, it’s a natural moment to rethink how data flows between them, and to finally deal with the old systems nobody dares to turn off.
- The data team becomes the bottleneck. When the backlog of data requests grows faster than the team can clear it, hiring more people won’t fix it. You need a foundation that lets people serve themselves.
When you probably don’t need one yet
I think many people skip this part, so I want to say it clearly. For many early-stage startups, the bigger risk is not the lack of a platform. It is spending months building one, while the team should focus on building the product. Of course, this is different if data is the core of your product. In my opinion, you can wait if:
- You have only a few data sources, and they rarely need to be combined
- The tools you already have (your product’s database, your CRM’s built-in reports) answer the questions people actually ask
- Nobody would use the platform yet. A platform without real use cases behind it turns into an expensive, empty warehouse
- You can’t commit anyone to own it. A data platform isn’t a one-time project. Without an owner, it will slowly break down
If you are not there yet, start small: put a BI tool on a replica of your database, write down how your key metrics are defined, and revisit the question in six months. And when the time comes, the first version doesn’t need to be big. A cloud data warehouse, a ready-made ingestion tool, a BI tool, and one person who is responsible for the data, are already enough for quite a long time.
One platform or several?
After you decide that you need a platform, there is one more related question, and it mostly matters for larger companies: does the whole company need one, or does each part of the business need its own?
One of the jobs of a data platform is to tie together data that’s scattered around the company, so it’s naturally centralized. For small and mid-sized companies, and for big companies with a single line of business or well-established processes, one central platform can hold all of the company’s data. It connects data from every system and every team, and keeps different teams from building the same thing twice.
But if a company has lots of brands or a really complex business, the teams may not need to connect their data that much, and a single central platform just isn’t worth as much. For example, if a company has several brands that run independently, or business lines in totally different industries, it may make more sense for each of them to build its own platform. At large companies, I’ve also seen a single central data team turn into a bottleneck when dozens of departments all need things from it at once. And at that scale, the hardest part is often not the technology. It is to get many teams to agree on the same definitions, and on who owns what. In short, it depends on how your company is structured, how big it is, and whether connecting data across teams actually creates more value.
Even with several data platforms in one company, you can still keep maintenance costs down by sharing the same cloud infrastructure and picking similar tech stacks. And having several platforms doesn’t mean the data is split up forever. If they need to work together later, there are ways to do it. Data federation uses a query engine that can read from several platforms in one query, so people can get data from all of them without copying everything into one place. Data mesh is more about organization: each business domain owns its data and publishes it as a product, following shared standards, so other teams can find and use it easily.
If you decide to build one
How to build a data platform is a big topic, and I will write another post about it. But here are a few things I would keep in mind from the first day, no matter how big the company is:
- Start from use cases, not tools. Write down who will use the platform and for what before you pick any technology. Build for your real business needs, not for showing off to other engineers. If your analysts look at yesterday’s data every morning, you don’t need real-time streaming.
- Go cloud by default. For most companies, especially small and mid-sized ones, cloud services mean no hardware to buy, scaling up and down as needed, paying for what you use, and built-in tools that help with compliance. But keep in mind that compliance is still your own responsibility, the cloud provider only covers its part. If you have strict security requirements, hybrid setups let you keep sensitive data on-prem and still use the cloud for the rest.
- Keep the raw data, and model the rest. Keep data close to its original form so you can always rebuild, then build clean, well-defined tables on top for people to use.
- Make trust a feature. Check data before it reaches a dashboard, and make sure jobs can be rerun safely. One wrong number in a meeting with executives can destroy the trust you built in months.
- Give it an owner. A platform needs a team responsible for it, and people on that team who understand the data itself, not just the code.
Wrapping up
In my opinion, a data platform is not something a company needs from the first day, and it is not something to build just because other companies have one. It becomes valuable when the data needs start to cross teams and systems, when people can’t agree on numbers, when analysts spend their days cleaning data, and when regulations or ML projects need data you can trust.
So the real question is not “should we have a data platform?” It is “how much are our data problems already costing us?” Once this cost is bigger than the cost of building and running a platform, it is the right time.