Sunday, 26 July 2015

The Data Science Spectrum

In order to describe what is data science, it has become increasingly popular to depict common data science skills as Venn diagrams, typically using three sets. Examples of this can be found in different sources such as IBM or Datajobs.com. However, I think the most popular Venn diagram is the one that Drew Conway included  in his site.

Some of these diagrams are used to describe how difficult (or impossible!) is to find a person that covers all the data science skills, to be placed in the centre. That would be a Unicorn, as MediaShift.org describes it. Other diagrams however are built just to describe what a data scientist precisely is: That person that contains all those features and is placed exactly at the centre, with the right balance of everything. 

The salad of data

Although it would be possible to meet all the skills and requirements, I personally think that seeking that type of person is slightly optimistic and perhaps misleading. An employer that needs a kind of handyman that would complete all the data science tasks would be similar to searching for a single person to plant, grow, harvest, cook and serve all the necessary vegetables for us to eat a salad. It is likely that in some of these steps that person is weaker than in others, affecting the whole process. Therefore, the final salad we obtain wouldn't be an optimal output. Following the salad metaphor, we can divide the steps in this way:

1) The planting and growing steps would refer to all those Extraction, Transformation and Load (ETL) tasks that must be done in order to retrieve data streams from the cloud or various databases and map all that information. Those ETL processes are a must in data science, like the job of the farmer if we want to obtain a salad.

2) It would be great if the harvesting and cooking of data could be always implemented during the ETL development. However, in many cases the data we have to deal with will be complex and therefore it would be very time and resource consuming to develop an ETL software able to provide every possible view of the data we might want in every specific case. Hence, in many cases it is necessary to add a layer of intelligence in order to retrieve the figures we might need for a specific report or combine several outputs to draw particular conclusions. That job can be done by an statistician or a pure scientist: a person able to remove noise, examine the distribution of data and aggregate information in order to make sense of it or to design and develop automated intelligent processes - that is, Machine Learning - to obtain these analyses in the future. Such processes are as important as the previous ones: we need a chef with the expertise to select the best veggies from the farm and cook them appropriately!

3) Finally we need to serve that salad in a nice way if we want somebody to enjoy eating it. That would be the job of the designer, able to arrange the findings in the data, present and sell them in such an appealing way that no senior manager or customer would be able to resist our tempting data salad!

I believe we could agree that the farmer and the chef might share some skills: they know the raw product and how to deal with it. Likewise, the chef and the designer know the final output (how it needs to taste and what ingredients should contain). However, I struggle to find the relationship between the designer and the farmer. It is not very clear to me what the Venn diagrams really mean with the intersection of these two. And this is why in my head I see data science not as a Venn diagram, but as an spectrum

The Data Science Spectrum

Although we should bear in mind that the salad metaphor is a bit simplistic when it depicts data science as a linear process - when in reality is not most of the time -, it is also true that not all the skills required for data science have to intersect, as the Venn diagrams imply. The spectrum alternative would look more like this:.



The Data Science Spectrum

Any data scientist should have the necessary skills to manipulate and cook data. That is, she/he should be a chef able to at least design Machine Learning, Artificial Intelligence and/or Statistic analyses on the data, as aforementioned. Therefore, any data scientist should touch the yellowish part of the spectrum in some way or another. In fact, I can understand that many data scientists would argue that their job is just that. However, since we almost never work in isolation, software development and/or communication skills are an essential part of most roles, so having some of the skills of the farmer or the designer would definitely give us a competitive advantage over other chef-only candidates when applying for a data science position.

In this regard, the more a data scientist stretches her/his skills towards the edges of the spectrum, the more valuable she/he would be in most data science jobs advertised nowadays.

By expanding the skill set to the left side of the spectrum, we would be able to automate processes once the results (of a new Machine Learning algorithm or a statistical process) are valid so they can be replicated when new data comes without any extra effort. Also, the data scientist would be able to design and develop ETL software to obtain and clean further data to improve any analysis. In this regard, the data scientist would become a Data Scientist - Developer, able to work independently without having to wait for other developers to complete any programming.

On the other hand, by stretching the skillset towards the right end of the spectrum, the data scientist' role increasingly becomes that of a Story-Teller, able to make sense of the results from a business perspective and compile crucial findings into infographics, presentations and other reports. This type of data scientist is becoming more and more popular in the UK, specially for statisticians or scientists able to use R or Matlab but lack coding skills.

It is worth mentioning that while the left side of the spectrum is focused on automatising processes, the right one performs the exact opposite: Generate unique outputs for specific cases.

I think that none of the pure milestones inside the data science spectrum would be considered data science, in the classic definition. There are other job titles that already cover those cases: Database and data warehouse developers/engineers, scientists/statisticians and business intelligence (BI)/public relationship professionals, respectively. However, the guys located between the milestones are more likely to be data scientists in the definition that I believe the buzzword refers to: A person that can plant, grow and harvest vegetables very well, or one that cooks and/or serves them in a highly professional way. That is, a Data Scientist - Developer or a Story-Teller.

It is perfectly fine for any data scientist to try to be that perfect unicorn placed exactly at the bullseye of the Venn diagrams or attempts to cover the entire spectrum. Having a job that requires us to complete all those different tasks would be fun and never monotonic! That would be great, but I wonder if it wouldn't be better to be part of a team where one or more members master some of the skills, complementing and learning from each other. In that regard, the team itself would be that unicorn, from day one. I think that is far easier and faster to achieve, isn't it?

Many would argue that a small/medium size company wouldn't have the resources to invest in such big data science team, but I think that having such a team should not imply any extra cost in the data science budget, at least compared to having one overflown, exhausted unicorn!

In some future post, before or after exploring the Data Scientist-Developer and Story-Teller, I will try to come back and dig into this and explain how the financial maths for this team would work without having to add extra budget.

As I mentioned in the first entry of this blog, I would love to discuss with any of you, keen readers your views and fix anything that might be wrong, so comment is free!

Thursday, 23 July 2015

Data Science, or whatever they call it

Coming from an academic background, appending the word "science" to such a generic term as "data" could be considered almost an insult. If you think about it, it seems a meaningless concept since everything in the world can be measured today. Even for the things that apparently can't be measured like happiness or love there is a gazillion of metrics proposed to compare different countries, individuals within societies or for psychological purposes, for example. Hence everything is data, and the science of data cannot be another thing but the science of everything (!), which sounds a little bit generic, to say the less.

However, it is easy to recognise that both ideas, data and science, work quite well as a buzz word, so I guess that is why nowadays we see data science almost everywhere and in every possible field, ranging from health to sport and from finance to ecology. Therefore, it looks like in its generality is actually its greatest virtue as a concept; but also it is its greatest inconvenience, because very few people in this planet would be able to tell you with a simple sentence what a data scientist is or what she/he does. Even many data scientists and companies trying to hire data scientists usually struggle to clearly define what they expect to do or to be done in a data scientist role. Many times the term data engineer appears as an attempt to differentiate the guys who are supposed to design/implement ETL processes and dig into endless log files (those regex heroes), from those who write equations impossible to understand to the common citizen and from those who collect whatever ugly output and makes it up into a gorgeous infographic.

You could find a myriad of articles here and there in newspapers, blogs and in recruiters' websites. Usually vaguely written with few guidelines to follow. And this is precisely what I consider to be the main challenge for this blog: to provide a wide description of the roles, persons and skills, with some dives into the deeps of the technicalities of the data science ocean. A broad user guide for the intrepid student who wants to climb the career ladder from the data face of the mountain, or for the confused manager who just wants something to be done and needs the right guy who knows how to use the right tool for the job. 

Obviously, that is something that I only aim to describe from my own experience, since it would be rather presumptuous even to attempt to achieve such a task by oneself. In this regard every comment, suggestion, feedback, guide, critique or correction, constructive or otherwise, would be greatly appreciated, howdy visitors!

It subsequent entries I will try to define this amazing field of study by describing the experiences I acquired over the years I have been travelling inside this world, in academia and in industry, from the more general overview to the problems different data scientists solve in their everyday workload and the common skills that would define what a data scientist is and why she/he is different from a statistician, a developer or a graphic designer. That is, to draw the big picture of what a data scientist is, or whatever they call the person who does that job.

Welcome to yet another data science blog!