Jasper Alblas
Jasper Alblas
Data Engineering • Analytics • Business Intelligence
You’re on Part 0.3 of 5 in Foundations → the hands-on build (QuakeFlow) starts at Part 1.1
If you’re new to data engineering, it’s tempting to jump straight into learning Python or SQL — in fact, that’s exactly what most beginners do. The problem is that without understanding how the data world fits together, the tools quickly become confusing. This series is built differently: understand why the tools exist and where they fit, before building a complete end-to-end project.
| Part | Topic |
|---|---|
| 0.1 | How the Data World Works |
| 0.2 | Data Careers Explained |
| 0.3 | Is Data Engineering Right for You? (This article) |
| 0.4 | Understanding Data Storage |
| 0.5 | Understanding Data Pipelines |
| 1.1 | Setting Up Your Data Engineering Environment — QuakeFlow begins |
| 1.2 | Your First Data Source: Working With a CSV |
In Part 0.1, you learned that data usually moves through a lifecycle: Source Systems → Ingestion → Storage → Transformation → Modeling → Serving → Consumption. In Part 0.2, you learned the main roles that work on each part of it: Data Engineer, Analytics Engineer, BI Developer, Data Analyst, Data Scientist.
Now it’s time to narrow the focus to one question.
DIs Data Engineering a good career path — and is it right for you specifically? It’s a strong path — technical, practical, and increasingly important as companies lean more on data, analytics, and AI. But it isn’t the right path for everyone. Some people enjoy building systems behind the scenes; others prefer analyzing data, building dashboards, presenting insights, or working with machine learning models directly.
The point of this article isn’t to convince everyone to become a Data Engineer — it’s to help you make a better decision, before you invest months learning it. By the end, you’ll know what Data Engineers actually do, what the work feels like day to day, what skills matter most, what kind of person tends to enjoy it, what the real downsides are, and whether you should continue into the hands-on part of this series.ore practical.
A Data Engineer builds and maintains the systems that make data available, reliable, and usable — typically working with pipelines, databases, warehouses, lakes, APIs, files, SQL transformations, scheduling tools, monitoring, and data quality checks.
Data Engineers build the infrastructure that moves data from source systems to analytical systems.
Picture a team that wants a dashboard tracking earthquake activity by region. The data needed for that — event readings, station metadata, regional groupings, timestamps — usually lives scattered across several systems. A Data Engineer collects it, cleans it, organizes it, and makes it available for reporting. That’s not just copying data from A to B — good Data Engineering means making sure the data is correct, complete, timely, secure, understandable, maintainable, and scalable. If a dashboard shows the wrong numbers, the business may make bad decisions. If a pipeline fails, reports don’t refresh. If a table is poorly designed, every analyst downstream inherits the pain. This is foundation work — done well, it’s what lets everyone else work confidently with data.
A beginner might imagine the job as writing code all day. There’s coding, but the work is broader. A realistic day might look like:
09:00 — Check whether last night’s data pipelines succeeded.
09:30 — Investigate why the earthquakes table has fewer rows than expected.
10:30 — Write SQL to clean and standardize station metadata.
11:30 — Join a meeting with an analyst to understand a new reporting requirement.
13:00 — Update a Python script that pulls data from the USGS API.
14:30 — Review a colleague’s pull request.
15:30 — Add a data quality check to catch missing magnitude values.
16:30 — Document a new table so other people know how to use it.
Some days are focused on building new pipelines. Some are focused on debugging. Some are focused on performance. Some are focused on understanding business logic. This mix matters — Data Engineering isn’t only technical, it also requires communication and judgment about how data is actually used.
Data Engineers often solve problems like: why did the pipeline fail last night? Why are there duplicate event IDs? Why doesn’t this region’s count match what Finance-adjacent reporting shows? How do we collect data from this new API cleanly? How should these tables be structured? How do we make this query faster? How do we detect bad data before users see it? How do we recover cleanly if something fails? How do we scale this as data volume grows?
Notice how few of these are “write this algorithm” problems. They’re systems questions — about how data moves, changes, breaks, and gets used. That’s why the role rewards people who enjoy structured problem-solving.
ItIt’s also useful to be clear about what’s outside the role. Data Engineering usually isn’t mainly about designing beautiful dashboards, presenting insights to executives, running statistical experiments, building machine learning models, or writing business strategy recommendations. Those tasks belong more to BI Developers, Data Analysts, or Data Scientists — Data Engineers may work closely with them, but the core responsibility is different.
A useful mental model: Data Engineers make data available. Analytics Engineers make data consistent. BI Developers make data accessible. Data Analysts make data understandable. Data Scientists make data predictive. These boundaries blur constantly in real companies, but the model is still worth having in your head.
You don’t need to master every tool before starting — but a few core skills matter again and again.
Essential. Data Engineers use SQL to query, validate, transform, and join data, create tables, investigate issues, and test assumptions. Even alongside Python, Spark, or cloud tools, SQL remains one of the most-used skills in the role — and it’s one of the best places for a beginner to start.
Widely used for reading files, calling APIs, automating tasks, cleaning data, moving data between systems, and writing pipeline logic. You don’t need to become a software engineer overnight — you need enough Python to write clear, reliable scripts. It’s one of the first practical tools you’ll use in the hands-on part of this series.
Structuring data so it can actually be used — fact tables, dimension tables, primary keys, foreign keys, grain, relationships, star schemas. Good modeling makes reporting easier; bad modeling makes everything downstream harder. Later in this series, you’ll build a real dimensional model for the QuakeFlow project.
Pipelines, SQL scripts, and configuration files should never live only on one person’s laptop with no history. Git helps you track changes, collaborate, review code, roll back mistakes, and build a portfolio you can actually show an employer — which is exactly why it appears before the hands-on project begins.
One of the most important skills in the role. Pipelines fail. APIs change. Files arrive late. Column names shift. Source systems contain bad data. Queries get slow. A large part of the job is asking: what broke, where did it break, and how do I fix it properly? If solving that kind of puzzle sounds satisfying, that’s a good sign. If technical troubleshooting sounds draining, the role may become frustrating.
Modern data platforms often run on Azure, AWS, or Google Cloud. You don’t need to be a cloud architect as a beginner, but understanding storage, compute, permissions, databases, basic networking, and cost awareness helps — and becomes more important as you move from beginner projects to production systems.
Often underestimated. Data Engineers regularly talk to analysts, BI Developers, product teams, data scientists, software engineers, and managers — and need to ask things like: what actually counts as a “significant” earthquake for alerting purposes? Should preliminary, unreviewed readings be included? Which timestamp defines when an event “happened” across time zones? Technical skill matters, but an unclear business definition can break a data solution just as thoroughly as bad code.
You like building things. Data Engineering is a builder role — pipelines, tables, processes, systems. The result isn’t always visible to end users, but it enables a lot of value downstream. If you enjoy creating something reliable that others rely on, that’s a good sign.
You enjoy debugging. The dashboard didn’t refresh. The source file is missing. The row count changed unexpectedly. The API returned a new field. A query that used to run in two minutes now takes forty. If investigating problems like these sounds interesting rather than painful, this role likely fits well.
You think in systems. A change in one place can ripple into failures several steps downstream: source system change → pipeline breaks → warehouse table missing data → Power BI dataset fails → business report is wrong. You need to think beyond the single script or table in front of you, to the whole flow.
You like structure. Naming conventions, folder organization, table design, documentation, testing, repeatable processes, clear ownership. Messy data systems get expensive fast — if you naturally like organizing things clearly, that’s genuinely useful here.
You’re comfortable working behind the scenes. When a dashboard looks good, the BI Developer often gets the praise. When the numbers support a decision, the analyst presents the insight. The Data Engineer made all of it possible — but usually from behind the scenes. If you need to be the one presenting the final insight, Data Analysis or BI may suit you better. If you like enabling others through strong foundations, this can be genuinely satisfying.
Data Engineering may not be the best fit if you strongly prefer presenting insights, working mostly with stakeholders, designing dashboards, avoiding code or debugging, working mainly in Excel, focusing on statistics or machine learning, or short, visual feedback loops. That doesn’t mean you can’t learn it — it just means another role may feel more natural.
If you enjoy business questions, consider Data Analysis. If you enjoy dashboards and semantic models, consider BI Development. If you enjoy SQL transformations and business definitions, consider Analytics Engineering. If you enjoy statistics and prediction, consider Data Science. The best role isn’t the one with the most impressive title — it’s the one where the daily work matches your strengths.
No career path is perfect, and pretending otherwise would defeat the point of this article.
1. Pipelines break. Source systems change, APIs go down, files arrive late, permissions expire, someone renames a column. Good teams reduce failures with testing, monitoring, and clear processes — but failures never disappear completely.
2. The work can be invisible. When everything works, people may not notice. When something breaks, everyone notices immediately. The better your work is, the less visible it may become — frustrating if you want frequent recognition.
3. Requirements are often unclear. “Can we get earthquake counts by region?” sounds simple until you’re untangling what counts as a “significant” event, whether preliminary readings should be included, which timestamp defines an event when it’s near a time-zone boundary, or how to handle an earthquake occurring right at a regional border. A large part of the job is surfacing hidden assumptions before they become bugs.
4. Tooling changes quickly. Tools come and go, best practices evolve, cloud platforms add new services constantly. The good news: fundamentals change far more slowly than tools — which is exactly why this series starts with concepts before technology.
5. Junior roles can be competitive. Data Engineering is in demand, but companies often prefer some evidence of ability, since the work touches production systems. This isn’t a wall against beginners — it’s why a real portfolio project matters, and exactly why this series is built around one instead of theory alone. project instead of only theory.
Data Engineering has grown more important as companies depend on data for reporting, analytics, machine learning, AI applications, automation, compliance, and product decisions. The more companies invest in analytics and AI, the more they need reliable data foundations underneath it. AI doesn’t remove the need for Data Engineering — in many cases it increases it. Models need clean training data. Dashboards need trusted metrics. Applications need reliable pipelines. Data quality matters even more once automated systems start depending on it. That’s why this remains a strong long-term career path.
Salaries vary by country, industry, seniority, and company size, so treat any figure as a rough signal rather than a promise. A broad European range might look something like: Junior Data Engineer, roughly €40,000–€65,000; Mid-Level, roughly €65,000–€90,000; Senior, €90,000 and upward, often significantly more at larger tech companies. As skills grow, common next steps include Senior Data Engineer, Analytics Engineer, Data Platform Engineer, Cloud Data Engineer, Data Architect, Analytics Architect, or Engineering Lead.
Yes. A relevant degree in Computer Science, Engineering, or Mathematics can help, but it isn’t the only path. Employers mostly care about whether you can do the work — write SQL, use Python, structure a project, work with data files, build a basic pipeline, use Git, understand databases, debug problems, and explain your design choices. For career changers especially, a finished portfolio project demonstrates practical ability far better than saying you completed a course — which is exactly why the hands-on part of this series is built around a realistic project.
Depends heavily on your starting point. A realistic beginner path might look like: SQL, basic Python, and Git in months 1–2; data files, APIs, databases, and simple pipelines in months 3–4; data modeling, warehouse concepts, and Power BI basics in months 5–6; transformation tooling, orchestration, testing, and a portfolio project in months 7–9; interview preparation and deeper cloud skills in months 10–12. Not a strict timeline — some people move faster, some slower. Consistency matters more than speed: one focused hour daily beats one intense weekend a month.
You may enjoy Data Engineering if you answer “yes” to many of these: Do you enjoy solving technical problems? Do you like understanding how systems connect? Are you willing to learn SQL and Python? Do you enjoy debugging? Do you like making messy things organized? Are you comfortable working behind the scenes? Do you care about reliability and correctness? Can you tolerate unclear requirements? Are you interested in databases and pipelines? Do you enjoy building foundations that others use?
You may prefer another path if you answer “yes” to many of these instead: Do you mainly want to present insights? Do you dislike coding? Do you prefer visual dashboard work? Do you want to focus mostly on business strategy? Do you dislike debugging technical problems? Do you prefer statistics and machine learning over pipelines? Do you want frequent, visible recognition from business users?
There’s no perfect score here — the purpose is noticing which kind of work sounds energizing and which sounds draining.
Continue if you want to learn how to build a practical data pipeline from raw data to business reporting. The hands-on project will give you real experience with project structure, Git, Python ingestion, data cleaning, SQL transformations, dimensional modeling, data quality thinking, Power BI reporting, documentation, and end-to-end data flow.
You don’t need to know everything before continuing — that’s the entire point of the series. But you should be willing to actually practice. Reading about Data Engineering helps; building something teaches far more.
The actual hands-on project you’ll build starting in Part 1.1 is QuakeFlow, using real earthquake data from the USGS catalog — not a fictional company. It follows the same lifecycle introduced earlier:
Raw Data → Python Ingestion → Staging Area → SQL Transformations → Dimensional Model → Power BI Dashboard
Working through it will show you directly whether you enjoy moving data, cleaning it, structuring tables, debugging real problems, thinking in pipelines, and connecting backend work to business reporting. If you enjoy the project, that’s a strong signal Data Engineering or Analytics Engineering suits you. If you find yourself far more drawn to the Power BI and analysis end, BI Development or Data Analysis may be the better fit.
Either outcome is a good outcome. The goal isn’t just finishing the project — it’s learning what kind of data work you actually enjoy.
Do I need to be great at programming? No, not at the start. Focus first on writing simple, readable Python scripts — you don’t need advanced algorithms to begin.
Do I need advanced mathematics? Usually no. Data Engineering sits closer to software engineering and databases than to statistics. Math matters more if you later move toward Data Science or Machine Learning.
Is SQL or Python more important? Both matter, but SQL is usually the stronger first skill since it appears in almost every data role. Python becomes especially important for automation, ingestion, APIs, and pipeline logic.
Do I need cloud skills immediately? No. Understanding files, databases, SQL, Python, and pipelines locally first will transfer directly once you move to cloud services later.
Can BI experience help me become a Data Engineer? Yes, often significantly. If you already understand reporting, data models, and business definitions, you understand the downstream half of the job — you’ll mainly need to build upstream skills: Python, pipelines, storage, Git, orchestration, and cloud basics.
Is Data Engineering better than Data Analysis? Not better — different. Data Engineering is more technical and systems-focused; Data Analysis is more business-facing and insight-focused. The right choice depends on which kind of work you enjoy.
If you’re still unsure, don’t decide from reading alone — build something. Even a small project teaches you more about your own preferences than any amount of research. You might discover you love writing pipelines. You might discover you prefer modeling data in SQL. You might discover dashboards are the part you enjoy most. You might discover debugging pipelines isn’t for you. All of those are useful discoveries. The best way to choose a data career isn’t picking the best-sounding title — it’s trying the work. That’s exactly what this series is built to help you do.
You now understand how the data world works, which data roles exist, and whether Data Engineering might fit you. The next step is preparation — before the hands-on project begins, we need to cover a bit more theory.
Next up → Part 0.4: Understanding Data Storage — learn about all the ways data can be stored.
[…] Is Data Engineering Right for You? […]
[…] Is Data Engineering Right for You? […]
[…] Is Data Engineering Right for You? […]
[…] Is Data Engineering Right for You? […]