A while ago I made a blog and vlog about the difference between data lakes and data warehouses. One of our most read and watched posts. Time for an update - partly due to the rapid growth of variants such as 'data lakehouses'. By Harmen, CTO & senior data scientist at Datalab
How did it work again…?
Data warehouse, data lake, data lakehouse… how did it work again? Data warehouses have actually existed for decades. Since the 2010s, these types of data storage locations have truly become commonplace. Very functional, but only when the prerequisites are right. One of the most important prerequisites is the storage structure. There are different styles of how you structure information: as fact-and-dimension tables (a so-called Kimball schema, named after its inventor) or close to the original dataset (often a normalised model). This latter model gives the most flexibility - the reason we are fans. You simply create derived datasets - so-called views - in which you shape the data precisely for your dashboard or analysis. If in the future you want to centre a different perspective on your data (for example from marketing to procurement logistics), you only need to create a new view. Creating views is straightforward: with our refined training, we teach our clients to create views in just a few hours. The time investment is therefore in storing data in a well-structured manner.
The shared network drive
Data lakes probably exist much longer than data warehouses, albeit not under this name. The precise definition of a data lake is vague, but in any case you store information without too much structure in an easily accessible location.
Previously, organisations had the shared network drive for this (or whichever drive letter - in any case a shared network drive). That turned out to be somewhat chaotic and access management is difficult to arrange. With the rise of the cloud, the data lake was born. Result: a network drive in the cloud with often unfindable information, unclear ownership, unclear structure, unclear which (parts of) data have been updated, and so on. At the same time, it is cheap and a quickly arranged place where data comes together. With the enormous flight that 'big data analytics' took around 2015, a convenient way for many organisations to get started. The technology has unfortunately not stood the test of time, as many organisations choose to migrate to a data warehouse - precisely because information is so hard to find. Data lakes are now primarily used for storing large datasets that are inherently unstructured, such as photos, videos, and backups. They also regularly serve as a temporary 'parking spot' for datasets that are later processed into a data warehouse.
Enter: data lakehouses - the ideal compromise between warehouse and lake?
A data lakehouse is a combination of a data warehouse (structured) and data lakes (unstructured). The trick is that data from both types of sources is easily combinable. In the lakehouse itself, you maintain metadata describing what is in the individual files in your data lake (and who is responsible, who may access it, and so on). That seems ideal: less effort to structure data, yet accessible and linkable, with good metadata.
For some organisations, a data lakehouse will indeed work fine - for example, organisations that work with frequently changing data sources and have their own data science department. The bigger players, in other words. Why particularly for them? Because data scientists do not need to wait for data warehouse experts to integrate data. This way, analysts can get to work faster.
"Ne sutor ultra crepidam"
Naturally… there is a downside to the data lakehouse. Summarised, it is best described with a well-known saying: "cobbler, stick to your last." Most data scientists and dashboard builders are handy with data, like a technically trained person is handy with cars. Even if you understand how a car works and can change the spark plugs, oil filters and oil yourself, that does not make you a car mechanic. It could well be that you use the wrong plugs or oil, or do other impractical things that in the longer term cause more damage than good.
More practically for the data analyst: there is quite a lot involved in properly structuring data. It is not without reason that renowned research firm Gartner writes that the data engineer role is indispensable in many teams.
Properly handling data takes time regardless…
With a data warehouse, you invest time upfront in structuring your data. With a data lake, far too often no time is spent on structuring data at all. And with a data lakehouse, you spend time structuring your data every time you want to create a new analysis.
Therefore: a data lakehouse can be a useful solution, especially for larger teams where data reuse is less of a concern and there are frequently changing data sources. Here too, existing technologies exist for a reason: because they work well. That certainly applies to a data warehouse. Whether a data lakehouse is also a good solution in the long term remains very much the question. Regardless, it is important to follow developments closely - something we at Datalab are happy to do.
Want to know which technology best suits you? Feel free to send an email or schedule a conversation. I am also very curious about your experiences with data warehouses, lakes and lakehouses. Do share!
