In order to describe what is data science, it has become increasingly popular to depict common data science skills as Venn diagrams, typically using three sets. Examples of this can be found in different sources such as IBM or Datajobs.com. However, I think the most popular Venn diagram is the one that Drew Conway included in his site.
Some of these diagrams are used to describe how difficult (or impossible!) is to find a person that covers all the data science skills, to be placed in the centre. That would be a Unicorn, as MediaShift.org describes it. Other diagrams however are built just to describe what a data scientist precisely is: That person that contains all those features and is placed exactly at the centre, with the right balance of everything.
The salad of data
Although it would be possible to meet all the skills and requirements, I personally think that seeking that type of person is slightly optimistic and perhaps misleading. An employer that needs a kind of handyman that would complete all the data science tasks would be similar to searching for a single person to plant, grow, harvest, cook and serve all the necessary vegetables for us to eat a salad. It is likely that in some of these steps that person is weaker than in others, affecting the whole process. Therefore, the final salad we obtain wouldn't be an optimal output. Following the salad metaphor, we can divide the steps in this way:
1) The planting and growing steps would refer to all those Extraction, Transformation and Load (ETL) tasks that must be done in order to retrieve data streams from the cloud or various databases and map all that information. Those ETL processes are a must in data science, like the job of the farmer if we want to obtain a salad.
2) It would be great if the harvesting and cooking of data could be always implemented during the ETL development. However, in many cases the data we have to deal with will be complex and therefore it would be very time and resource consuming to develop an ETL software able to provide every possible view of the data we might want in every specific case. Hence, in many cases it is necessary to add a layer of intelligence in order to retrieve the figures we might need for a specific report or combine several outputs to draw particular conclusions. That job can be done by an statistician or a pure scientist: a person able to remove noise, examine the distribution of data and aggregate information in order to make sense of it or to design and develop automated intelligent processes - that is, Machine Learning - to obtain these analyses in the future. Such processes are as important as the previous ones: we need a chef with the expertise to select the best veggies from the farm and cook them appropriately!
3) Finally we need to serve that salad in a nice way if we want somebody to enjoy eating it. That would be the job of the designer, able to arrange the findings in the data, present and sell them in such an appealing way that no senior manager or customer would be able to resist our tempting data salad!
I believe we could agree that the farmer and the chef might share some skills: they know the raw product and how to deal with it. Likewise, the chef and the designer know the final output (how it needs to taste and what ingredients should contain). However, I struggle to find the relationship between the designer and the farmer. It is not very clear to me what the Venn diagrams really mean with the intersection of these two. And this is why in my head I see data science not as a Venn diagram, but as an spectrum
The Data Science Spectrum
Although we should bear in mind that the salad metaphor is a bit simplistic when it depicts data science as a linear process - when in reality is not most of the time -, it is also true that not all the skills required for data science have to intersect, as the Venn diagrams imply. The spectrum alternative would look more like this:.Any data scientist should have the necessary skills to manipulate and cook data. That is, she/he should be a chef able to at least design Machine Learning, Artificial Intelligence and/or Statistic analyses on the data, as aforementioned. Therefore, any data scientist should touch the yellowish part of the spectrum in some way or another. In fact, I can understand that many data scientists would argue that their job is just that. However, since we almost never work in isolation, software development and/or communication skills are an essential part of most roles, so having some of the skills of the farmer or the designer would definitely give us a competitive advantage over other chef-only candidates when applying for a data science position.
In this regard, the more a data scientist stretches her/his skills towards the edges of the spectrum, the more valuable she/he would be in most data science jobs advertised nowadays.
By expanding the skill set to the left side of the spectrum, we would be able to automate processes once the results (of a new Machine Learning algorithm or a statistical process) are valid so they can be replicated when new data comes without any extra effort. Also, the data scientist would be able to design and develop ETL software to obtain and clean further data to improve any analysis. In this regard, the data scientist would become a Data Scientist - Developer, able to work independently without having to wait for other developers to complete any programming.
On the other hand, by stretching the skillset towards the right end of the spectrum, the data scientist' role increasingly becomes that of a Story-Teller, able to make sense of the results from a business perspective and compile crucial findings into infographics, presentations and other reports. This type of data scientist is becoming more and more popular in the UK, specially for statisticians or scientists able to use R or Matlab but lack coding skills.
It is worth mentioning that while the left side of the spectrum is focused on automatising processes, the right one performs the exact opposite: Generate unique outputs for specific cases.
I think that none of the pure milestones inside the data science spectrum would be considered data science, in the classic definition. There are other job titles that already cover those cases: Database and data warehouse developers/engineers, scientists/statisticians and business intelligence (BI)/public relationship professionals, respectively. However, the guys located between the milestones are more likely to be data scientists in the definition that I believe the buzzword refers to: A person that can plant, grow and harvest vegetables very well, or one that cooks and/or serves them in a highly professional way. That is, a Data Scientist - Developer or a Story-Teller.
It is perfectly fine for any data scientist to try to be that perfect unicorn placed exactly at the bullseye of the Venn diagrams or attempts to cover the entire spectrum. Having a job that requires us to complete all those different tasks would be fun and never monotonic! That would be great, but I wonder if it wouldn't be better to be part of a team where one or more members master some of the skills, complementing and learning from each other. In that regard, the team itself would be that unicorn, from day one. I think that is far easier and faster to achieve, isn't it?
Many would argue that a small/medium size company wouldn't have the resources to invest in such big data science team, but I think that having such a team should not imply any extra cost in the data science budget, at least compared to having one overflown, exhausted unicorn!
Many would argue that a small/medium size company wouldn't have the resources to invest in such big data science team, but I think that having such a team should not imply any extra cost in the data science budget, at least compared to having one overflown, exhausted unicorn!
In some future post, before or after exploring the Data Scientist-Developer and Story-Teller, I will try to come back and dig into this and explain how the financial maths for this team would work without having to add extra budget.
As I mentioned in the first entry of this blog, I would love to discuss with any of you, keen readers your views and fix anything that might be wrong, so comment is free!
As I mentioned in the first entry of this blog, I would love to discuss with any of you, keen readers your views and fix anything that might be wrong, so comment is free!
