Guangyang Li's Blog

What Do Data Engineers Actually Do?

Data engineering is a branch of software engineering. As more companies need people to work on their data, this role gets more and more attention. In this post, I will explain what data engineers (I will call them DEs below) are responsible for, the different kinds of DE jobs in the industry, and answer some common questions from people who are thinking about moving into this role.

So what does a DE do? In one sentence, the core job is to solve the company’s problems around storing and processing data. Company data solutions moved from relational databases to distributed databases, and then to cloud data warehouses and data lakes. The job titles changed along the way too, from database administrator and ETL engineer to data engineer. Today, when people say DE, they basically mean a role that works on big data processing.

More specifically, the work of a DE falls into two big areas:

  1. Data platform: building and maintaining the data platform, so the company can store, transform and analyze data
  2. Data business logic: turning business logic and processes into data pipelines in code, and managing the whole lifecycle of the data on the platform

But companies are different in structure, size and the solutions they choose. So in a real job, you may do only part of the work above, or you may do both areas.

For example, at Facebook (now Meta) and Uber , DEs work almost only on the business side. Most of their time is spent on calculating business metrics and building processing flows inside a quite mature internal data platform. In other words, their main tool is SQL, and their core task is building data jobs with SQL. The platform itself is developed and maintained by other teams made of software engineers.

On the other hand, many small and mid-sized companies use public clouds like AWS or GCP, and they buy ready-made data platform services from the cloud provider. This makes the infrastructure part much easier. DEs at these companies can build a data platform like playing with Lego, by picking the cloud services that fit best. These engineers sit between DE and SDE (software development engineer). They may be responsible for the platform infrastructure, and also for the business logic of data processing.

Expedia , where I used to work, is somewhere in the middle. A team of SDEs maintains a set of open source data lake components . Then the data team of each business line deploys these components on AWS by themselves, and builds and maintains a data platform just for their own department.

Data platform

Building a data platform includes these kinds of work:

  • Planning a stable and scalable storage and processing architecture. Is the incoming data mostly batch or streaming? Does the output need to be real-time, or is a scheduled update enough? How fast is the data growing? The solution should be designed around these real needs and scenarios.
  • Picking a good data orchestration tool and building reliable data pipelines. The most common tool is Airflow .
  • Providing support and management tools for the platform, like data catalog, data lineage, metadata management, access control, monitoring and alerting. These tools can be built in-house, or bought from a vendor.
  • Providing data quality tools that monitor things like freshness, accuracy and completeness of the data.
  • Providing query engines and query services or interfaces, so downstream teams can get data in the way they need.
  • Automation tools for operations, like Terraform and Ansible , and cloud services for the same purpose, like AWS CloudFormation .
  • Choosing cloud services. Know what the cloud providers offer and what popular SaaS data services are on the market, and decide which ones are worth bringing into the company’s platform.

How big this work is depends a lot on the size of the company and its tech choices. If a company runs its own data center, it probably needs to build its own version of almost every part of the platform, and each part could be a whole team. But for a company that already moved to the cloud, a small team of a few people can cover all of it.

Data business logic

Once there is a data platform, DEs on the downstream teams can run real data jobs on it. Turning business needs into data engineering tasks includes work like:

  • Designing and building data metrics. Metrics usually come from analysts on the downstream teams. But if you understand how a metric is calculated and why the business needs it, it helps a lot when turning the requirement into code.
  • Data modeling. Design data models that fit the needs of both the input and the output.
  • Handling source data in all kinds of formats, and then cleaning, standardizing and transforming it after it lands in the platform.
  • Tuning jobs and queries. The most common one is SQL tuning, and the way to tune depends on the query engine. Tuning processing jobs also depends on the tools the platform uses, like Spark , Hive , Flink and Kafka .
  • Monitoring job status. For example, whether the upstream database is stable, whether the services in the platform are healthy, and whether the data quality is good.

One thing to note: in real work, there is no clear line between the platform side and the business side. In both big and small companies, a DE’s actual job can be any mix of the points above.


What I described above is quite general, and probably too rough for people who really want to move into this role. So below I try to answer some questions that people who want to switch may ask.

Does a data engineer need strong coding skills?

It mainly depends on what the team needs from this role. In short, it depends on whether the job is closer to the platform side or the business side.

If the team builds the parts of the data platform by itself, the coding skills they expect from a DE are basically the same as a software engineer. The main output of these engineers is a service or software that people inside the company use to work with data.

But most DE jobs on the market are on the business side. The coding bar is a bit lower than for software engineers, but the bar for processing data with SQL is much higher. Typical coding work is writing data processing jobs and data orchestration. Besides, DE jobs usually ask for experience with big data processing, because many projects depend on the team understanding distributed tools. This kind of experience is also closely related to how well you can read and write code.

What is on-call like for a data engineer? Is it easier than for an SDE?

For the platform side, the job is basically maintaining a set of data tools for internal users. So on-call means keeping these tools working and handling the issues users report. Since all users are internal, it is much easier to talk to them and test things than with external users. They also usually don’t send strange inputs to your software. So the on-call pressure is relatively low.

For DEs on the business side, the on-call work and pressure depend on where their data sits in the company. If a company has many business lines, there is usually a central data team in the upstream that manages the core data, and each business line downstream has its own team for its business data. As you can imagine, the core data team has much more pressure than the business teams. The upstream team has to hold itself to strict SLAs, like a strict SLA on data freshness, so the downstream data can be available on time. Their on-call is usually about keeping the data jobs running normally. When something goes wrong, they need to find a workaround quickly and backfill the data, so they don’t break the SLA.

For downstream business teams, the pressure depends on their end users, which means the data consumers. Take a BI team as an example. Its data serves internal BI analysis, like business analysts building visualizations in Tableau. These users usually don’t need real-time data, and they are more tolerant of data delays than external users. So I would say this is the team with the easiest on-call. But if the end users are external, for example a government regulator, then a data delay or missing data could lead to a fine or even losing a license. In that case the team needs enough people to make sure the data is available, and the on-call pressure is quite high for a B2B product.

What career paths does a data engineer have?

As a branch of software development, the career path of a DE is not very different from an SDE. For engineers who mostly build data platforms, they can go deeper into platform architecture and become a data platform architect, or move to SDE roles that are more on the infrastructure side. For engineers who mostly work on the business side, besides moving to data infra SDE roles too, they can go deeper into a specific use of data. For example, they can become a machine learning engineer, a data visualization engineer, or a data scientist.

If you want to move into management, a data platform team is a good place to start. As data-driven decision making becomes more popular, the data platform is becoming a core team in many companies. Compared with business teams, it is usually more stable, and its work often has a bigger impact.