Unclear data structure

An answer to sometimes even a simple question is not always easy to obtain. Precisely at those moments when urgency is required to get a particular overview or a concrete answer. This also applies - or perhaps especially applies - to large (public) organisations that from the outside appear to have more than sufficient resources to keep things in order internally. How is that possible? Often, unclear data structure is the root cause.

Eventually you do manage to formulate an answer, if you invest enough time and money. Sometimes even with the help of external parties. To arrange processes so that you can adequately answer such questions yourself, many organisations encounter similar problems: inconsistent data structure, multiple truths, an inconveniently set up foundation, or applications that are difficult to unlock.

Structured processes

Due to the pressure of daily work, people never get round to restructuring processes. Employees are so busy searching for answers to - theoretically - fairly simple questions that the bigger picture is lost sight of. The result: a vicious circle of expensive inefficient processes with the accompanying frustration. Whilst the solution is fairly simple: ensure that available data is stored in a normalised and readable way. And train employees to work with it. I am convinced that everyone can write a SQL query with which you can see, for example, which products sell well, which customer groups are most attractive, or how much revenue was achieved in a year. That data does need to be available and usable, of course.

Multiple truths

When data is hidden away in an application and is not (correctly) unlocked, the following scenario often occurs: a colleague exports data. The data is then processed - transformations or filters - to answer a specific question. The export is then often saved on their own laptop. Subsequently, another employee comes along with a slightly different question and sets to work with the same dataset from the application, but not the dataset that their colleague has on their own laptop. The result is two different truths that can even contradict each other.

Now this example is not that difficult to resolve. The two colleagues talk to each other and together discover what the other did and how this led to the different truths. Good, solved. But suppose these two employees have both been doing this for years without knowing it of each other? Because these are Excel files or CSVs, there is of course no backlog, and over the years nobody knows precisely when they did what. Both employees then continue working in their own way and it remains a mystery which of the two colleagues is right.

This may sound like an unlikely situation, but it does happen regularly. Taking it one step further - what happens when one of the two retires? Who then manages the document? How do we safeguard all the knowledge it contains? Nobody understands the document any more or knows how certain transformations or calculations were done. In other words: the whole thing starts again from scratch.

The solution

Even in the short time I have worked at Datalab, I have seen such examples come past multiple times. As time progresses and people continue working in the 'old familiar way', it becomes ever harder to arrive at a solution. Whilst the solution is relatively simple - namely a well-organised data warehouse. It may sound like a sales pitch - that is not my intention. But it would not surprise me if these kinds of processes are already taking place in your organisation. And that it is becoming harder by the day to resolve.