A conversation between Harmen and Koen
Harmen: 'A significant part of a data scientist's work is actually the work of a data engineer. What exactly is data engineering and why would you specifically want to hire someone for it? We ask Koen, our data engineer. Koen, what do you do daily as a data engineer?'
Koen: 'As a data engineer, I am responsible for managing and setting up data flows. This concerns unlocking various online and local data sources. Cleansing them is also part of it. Ensuring that data is stored in the correct format so that an analyst can immediately perform an analysis without further effort.'
Harmen: 'Many people don't think of data engineering when they hear data science, or data unlocking at all - they think more of building dashboards. How does this role differ? Do companies sometimes have a blind spot for the function of data engineering?'
Data engineering versus data science
Koen: 'Yes, I think many companies don't have that clearly on their radar. I think most analysts, or people who build dashboards - those who visualise data - definitely deal with data engineering. Ultimately you need to obtain the data in the correct format from a particular source. But the moment this happens on a large scale, you absolutely need a data engineer. He or she focuses entirely on that task. Because don't underestimate the number of data sources involved once you take a large-scale approach. It quickly becomes too much for the analyst who really just wants to focus on the dashboard or the analysis.'
Harmen: 'What you regularly hear is that an analyst spends 80 per cent of their time cleansing and linking data. Is this also typically something you can deploy a data engineer for?'
Koen: 'Absolutely, I think most of the work that analysts spend their time on is cleansing data. So it only makes sense for someone to focus specifically on that, so that all that kind of work no longer falls to the analyst. The analyst can then fully focus on what they were hired for: producing an analysis or a dashboard. So yes, definitely!'
Harmen: 'You say "setting up the process" - what exactly do you mean by that? I can imagine what cleansing data involves, but does it also concern linking data, for example?'
Koen: 'I think everything from obtaining the data to preparing it for the analyst can be fully set up as a process. So indeed, unlocking from various sources, then - if needed - cleansing that data. Also linking data, perhaps reformatting it so that everything is available in the correct way for the analyst. As a result, the analyst no longer has to deal with any of that.'
Harmen: 'So the work of a data engineer also involves the software setup so that data flows are permanently unlocked. Is that correct?'
Koen: 'Absolutely. You spend quite a lot of time initially to thoroughly understand the process and estimate what is needed. Once that is done, you ideally want to spend as little time on it as possible afterwards. The process ensures that all these steps run completely automatically, so that ultimately a ready-made data source can be accessed by the data analyst.'
Harmen: 'That means a data analyst is well able to unlock a data source ad hoc. For example, when a small extra dataset is needed one-off, and that a data engineer is primarily deployed for datasets that will be used frequently?'
Koen: 'Yes, I think it is handy if an analyst first explores what data is needed. So perhaps it is good to do an ad hoc import from a data source first. Then, once it is certain that an analysis or dashboard goes into production, the data engineer sets up the process so stably that the analyst no longer needs to worry about it and the analyses keep running constantly, or the dashboard constantly receives the correct data.'
Harmen: 'In the past, the work of a data engineer in my experience consisted mainly of programming. Is that still the case? And are there other aspects involved?'
The work of a data engineer
Koen: 'A data engineer's work still primarily consists of programming. You ultimately need to programme those processes correctly. From experience I know that nowadays more than just programming is involved. You need to understand the logic of the data - that takes quite some time. You also need to understand well how a data source is structured, so understand what data you are bringing in. That is very relevant for being able to set up processes properly. Regular sparring with the analyst to keep their requirements in mind is also important. After all, you need to know exactly what they ultimately want and what they need for the analyses or the dashboard. And not unimportantly, you also regularly liaise with the end client. You need to know which data sources they want to unlock, which data they want to obtain from where. So much more is involved. It is a total package that you deal with today.'
Harmen: 'You don't have a background as a programmer yourself. How did you end up in this role?'
Koen: 'That is indeed not a very standard route I have followed. I have a background in political science and was particularly interested in the statistics component. Moreover, I found programming particularly interesting. Subsequently, after my studies, I naturally rolled into the role of data analyst. Gradually I dealt more and more with preparing datasets - something that an analyst generally spends most of their time on. Out of interest, I increasingly ended up in preparing and setting up data processes. That is now essentially what I mainly focus on.'
Harmen: 'Suppose someone wants to become a data engineer and has a background as a business analyst, BI specialist or data scientist - what would you advise them?'
Advice and tooling
Koen: 'Personally, I have benefited most so far from focusing on properly mastering a programming language, primarily Python. You also need to be handy with Bash commands - there is a lot to gain there. Learning to work well with cloud tooling. Not every company works in the cloud yet, but you do see that much data engineering today takes place in the cloud. That often makes it considerably easier and more flexible for a data engineer to run the workload in the cloud. So focus on that, ensure you become familiar with it.'
Harmen: 'You speak about the cloud - you can go in any direction with it. Which tooling do you use most yourself, or find pleasant to work with? So what would you recommend?'
Koen: 'Lately I have been working a lot with Apache Airflow. I am actively developing my skills in it. Apache Airflow is a Python-based ETL tool, but it is more than that. You can incorporate the full workload into it. It can even be deployed for training models, which you can also process in a workflow. Because I am already familiar with Python, that helps me well. Especially my Python knowledge helps in translating into building pipelines. Moreover, you have the full flexibility of Python when building your pipelines. There are of course other great tools, such as Azure Data Factory, a strongly visually oriented ETL tool. It has nice functionalities, but what appeals to me most about Airflow is the flexibility and retaining the programming language when setting up and orchestrating pipelines.'
Harmen: 'Now you speak from the role of data engineer. If you had to advise a company, would you also recommend Apache Airflow? And why should a company choose it?'
Koen: 'That depends on a company's strategy: if it truly has a long-term vision and wants to seriously invest in good engineering capacity, then I would indeed advise investing in this kind of tooling. I am convinced that this will deliver much added value in the long run. However, it also depends on the skills of the employees. Should people not be very familiar with programming, or specifically Python, then it is of course also possible to opt for a more visual tool such as Azure Data Factory to set up pipelines in a more click-and-track manner.'
Harmen: 'Koen, thank you for this explanation. I hope our readers have a good picture of the function and relevance of data engineering. Should you have questions about this topic or about data-driven work in general, call me. We can help you set up your data flows, or help find the right data engineer.'
