<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Data Visualisations |</title><link>https://carlos-hugoblox.netlify.app/en/tags/data-visualisations/</link><atom:link href="https://carlos-hugoblox.netlify.app/en/tags/data-visualisations/index.xml" rel="self" type="application/rss+xml"/><description>Data Visualisations</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-GB</language><lastBuildDate>Sun, 01 Sep 2024 00:00:00 +0000</lastBuildDate><image><url>https://carlos-hugoblox.netlify.app/media/icon_hu_aa3341e371185529.png</url><title>Data Visualisations</title><link>https://carlos-hugoblox.netlify.app/en/tags/data-visualisations/</link></image><item><title>Can Digital Goods be Neutral? Evaluating OpenStreetMap’s equity through participatory data visualisation</title><link>https://carlos-hugoblox.netlify.app/en/projects/can-digital-goods-be-neutral/</link><pubDate>Sun, 01 Sep 2024 00:00:00 +0000</pubDate><guid>https://carlos-hugoblox.netlify.app/en/projects/can-digital-goods-be-neutral/</guid><description>&lt;h2 id="the-project"&gt;The project&lt;/h2&gt;
&lt;p&gt;Equity is a fundamental pillar of any digital innovation aiming to produce positive societal impacts that characterises any digital good.&lt;/p&gt;
&lt;p&gt;Digital goods have typically been framed as mere technological artefacts, by consciously putting aside any human consideration to pursue a neutral approach. Conversely, Neutrality has been portrayed as a positive aspiration that prevents any form of bias that could pervert the fulfilment of the digital goods’ mission. We contend that, as well-intentioned as this aspiration may be, this apparently neutral standpoint that ignores how these goods are governed, may be inadvertently producing and reproducing new types of oppression and colonialism.&lt;/p&gt;
&lt;p&gt;This is especially critical in projects whose community are extremely biased towards specific and hegemonic demographics, like the case of OSM. The election of OSM as study case is relevant and timely: OSM is the largest and more successful collaborative map of the world, and in February 2024, it was recognised as a global Digital Public Good by the UN-backed Digital Public Good Alliance. Like Wikipedia, OSM is based on principles of collaboration, openness and neutrality, and its data, contributed by a global community of volunteers, complements official data sources, and populates thousands of popular tools and services.&lt;/p&gt;
&lt;p&gt;Our project examines and evaluates how the principle of neutrality driving many digital goods promotes or hampers equity by implementing a participatory process with Geochicas, a collective of women and LGBT+ mappers.&lt;/p&gt;
&lt;p&gt;We will implement a transformative participatory process to evaluate and surface how minoritised demographics are involved, recognised, or excluded from data production and decision-making in OSM while empowering its participants. We will do so in two stages. First, we will identify “gender controversies” in OSM, this is, specific examples of existing or missing map features and the decisions leading to that which are problematic from a gendered lens. Second, we will combine lived experiences with expertise in data visualisation to co-design a tool to represent, quantify and communicate those gender controversies to enquiry on how neutrality is being operationalised in OSM.&lt;/p&gt;
&lt;p&gt;We expect our findings to be returned to OSM and inform potential transformation in OSM’s governance, database, and representation that are guided by equity principles. More broadly, we expect the findings to be adapted to other cases of digital goods and initiate similar transformations.&lt;/p&gt;
&lt;h2 id="team"&gt;Team&lt;/h2&gt;
&lt;p&gt;Dr
, (PI)&lt;br&gt;
Centre for Interdisciplinary Methodologies, University of Warwick&lt;/p&gt;
&lt;p&gt;Dr
, (Co-I)&lt;br&gt;
Centre for Interdisciplinary Methodologies, University of Warwick&lt;/p&gt;
&lt;p&gt;Dr &lt;strong&gt;Selene Yang&lt;/strong&gt;, (Advisory Board)&lt;br&gt;
Geochicas (Founder), Wikimedia Foundation, Research fellow at Stanford Center on Philanthropy and Civil Society&lt;/p&gt;
&lt;p&gt;Dr &lt;strong&gt;Jorge Leon-Casero&lt;/strong&gt;, (Advisory Board)&lt;br&gt;
Universidad de Zaragoza (Spain)&lt;/p&gt;
&lt;p&gt;Dr &lt;strong&gt;Rachel Palmen&lt;/strong&gt;, (Advisory Board)&lt;br&gt;
Internet Interdisciplinary Institute (IN3), Universitat Oberta de Catalunya&lt;/p&gt;
&lt;p&gt;Prof.
, (Advisory Board)&lt;br&gt;
Centre for Interdisciplinary Methodologies, University of Warwick&lt;/p&gt;
&lt;h2 id="funder"&gt;Funder&lt;/h2&gt;
&lt;p&gt;This project has been funded by
&lt;/p&gt;</description></item><item><title>Considerations on the importance of data and science in data science</title><link>https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/</link><pubDate>Sun, 05 Apr 2020 00:00:00 +0000</pubDate><guid>https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/</guid><description>&lt;!-- Tip: open with the why, then show results, code, and next steps. --&gt;
&lt;p&gt;I must confess: during these days of lockdown, I have been toying with dashboards and infographics about COVID-19 outbreak, just like everyone else.
with nice plots, some formatted tables and even some (basic) maps. It&amp;rsquo;s shiny and nice, you can spend some time playing with it turning layers on and off, hovering, zooming&amp;hellip; And, honestly, I have to admit that I am proud of some of the results I have achieved. &lt;strong&gt;And yet, it is flawed. Just like everyone else&amp;rsquo;s&lt;/strong&gt; (or almost). And yet, I will keep improving it, even though I&amp;rsquo;m afraid it will always be flawed and even though I acknowledge that it will never be a contribution to improve knowledge on the topic.&lt;/p&gt;
&lt;p&gt;Why, then, am I persisting on keeping working on it if I know I cannot change its fate? Admittedly, at some point, I asked myself that very question and I even considered quitting. Not only I didn&amp;rsquo;t want to lose my time (even in these days when we are locked down at home there are plenty of things we can do), but I didn&amp;rsquo;t want to contribute to generating noise, misinformation and even more dramatism about an already important drama. Because that&amp;rsquo;s what flawed graphics do. But in the end, I realised that &lt;strong&gt;working on a dashboard like that could be a great opportunity for learning by doing.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Many types of lessons can be learnt from it, like those related to technical skills, or those related to how data is gathered, visualized and analyzed. &lt;strong&gt;Today, when data and figures on COVID-19 are everywhere, I want to share some reflections on the science (or lack of it) in data science.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="1-get-the-right-data"&gt;1. Get the (right) data.&lt;/h2&gt;
&lt;p&gt;As obvious as it sounds, there is no data visualization nor data science without data to visualize or analyze. Therefore, the first thing that anyone who wants to make any type of data visualization or data analysis is to get the data. Second: we cannot use any type of data. We need to use &lt;em&gt;good data&lt;/em&gt;, and &lt;strong&gt;by &lt;em&gt;good&lt;/em&gt; I mean &lt;em&gt;reliable&lt;/em&gt;, &lt;em&gt;usable&lt;/em&gt; (in terms of licences, formats and structure), up-to-date and frequently &lt;em&gt;updated&lt;/em&gt;, and, hopefully, &lt;em&gt;official&lt;/em&gt; data that is &lt;em&gt;representative&lt;/em&gt; enough to explain the phenomenon we are studying.&lt;/strong&gt; While this is usually non-trivial, it is even more crucial if we are to explain a completely new phenomenon that it is happening as we speak and it does at a global scale like COVID-19.&lt;/p&gt;
&lt;p&gt;Usually, there are only two possible options&lt;sup id="fnref:1"&gt;&lt;a href="#fn:1" class="footnote-ref" role="doc-noteref"&gt;1&lt;/a&gt;&lt;/sup&gt;: either gather data by ourselves or rely on others&amp;rsquo; data. Whereas gathering our own data might be the best choice for some scenarios, in the case of COVID-19, it is unlikely that we are in a position to gather the kind of data that might be useful for us (in fact, even governments are struggling to do so, as we will see). Therefore, we are left to just one option. Of course, we cannot rely on some random person or institution, we need to rely on someone we can trust, like universities (because they tend to provide rigorous data), governments (because they provide official data) or organizations (like the
). But where do we get the data from?&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/WHO-dashboard.png"
alt="WHO&amp;rsquo;s COVID-19 dashboard (Screenshot from 02-04-2020)"&gt;&lt;figcaption&gt;
&lt;p&gt;WHO&amp;rsquo;s COVID-19 dashboard (Screenshot from 02-04-2020)&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;While most governments provide data in open licences that allow its use and reuse for any purpose, they usually fail to address another important issue: the file format and data structure. As a result, in countries like the UK or Spain, data is not ready to consume as it is, without manual work&lt;sup id="fnref:2"&gt;&lt;a href="#fn:2" class="footnote-ref" role="doc-noteref"&gt;2&lt;/a&gt;&lt;/sup&gt;. Others, such as WHO, provide nice dashboards with data, but they do not provide the raw data to be consumed by 3rd parties. The good news is that there are people and institutions who are working in releasing clean data with open licences, such as
,
that gathers data from different governments and WHO, and even
that fetches data from UJH and provides a nice data frame for developers and data scientists to use in their projects.&lt;/p&gt;
&lt;p&gt;As a result, it is no wonder that most infographics and dashboards worldwide rely on the same data sources. It seems a sensible decision: not only we get data which is ready to use, but we do it from trustful sources. And yet, as I will argue, it is because of that reason that most of them are wrong. But what could possibly go wrong?&lt;/p&gt;
&lt;h2 id="2-dont-take-data-too-seriously"&gt;2. Don&amp;rsquo;t take data too seriously&lt;/h2&gt;
&lt;p&gt;Now that we know where to get the data from, there is something we have to be aware of: by relying on data generated by others we have not solved the (main) problem of data gathering, we have simply transferred the responsibility to somebody else, but someone still has to deal with what we have been trying to avoid. And, surprise, not every country gathers the data in the same way.&lt;/p&gt;
&lt;p&gt;Take the case of the most basic and crucial question: &lt;strong&gt;how are the number of confirmed cases defined&lt;/strong&gt;. Since COVID-19&amp;rsquo;s symptoms are very similar to those of influenza and the only way to know if someone is infected by it is by testing positive in the tests&lt;sup id="fnref:3"&gt;&lt;a href="#fn:3" class="footnote-ref" role="doc-noteref"&gt;3&lt;/a&gt;&lt;/sup&gt;. This seems a great definition: we have an objective test which is same the for everyone and all countries seem to use the same criteria. Unfortunately, those tests require equipment which is scarce (compared to the current worldwide demand), can only be made in hospitals, and require up to two days to get the results. Therefore, there are many other scenarios that are not considered within this test, such as those shown in figure 2a. &lt;strong&gt;So yes, every government provides that figure, yet all of them are much lower than the real figure. How much lower? There is no way to know.&lt;/strong&gt;&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/iceberg.png"
alt="Confirmed figures ≠ actual figures. There are many more cases than the official ones. How many more? We can&amp;rsquo;t possibly know"&gt;&lt;figcaption&gt;
&lt;p&gt;Confirmed figures ≠ actual figures. There are many more cases than the official ones. How many more? We can&amp;rsquo;t possibly know&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Let&amp;rsquo;s focus on another example: &lt;strong&gt;the number of deaths by COVID 19&lt;/strong&gt;. Apparently, this should be easier. Every country keeps a record on the number of deaths per day and its cause of death, so it should be easy to filter those who died from COVID-19 amongst all the possible causes. Well, no. As we have seen, if we can&amp;rsquo;t define with precision the number of people who are infected by COVID-19, we will not be able to know the number of people who have died as a result of it.&lt;/p&gt;
&lt;p&gt;But it can be even trickier if we consider that every country has different criteria on how they count the number of deaths of people who were tested positive&lt;sup id="fnref:4"&gt;&lt;a href="#fn:4" class="footnote-ref" role="doc-noteref"&gt;4&lt;/a&gt;&lt;/sup&gt;. Take the case of UK&amp;rsquo;s definition:&lt;/p&gt;
&lt;blockquote class="border-l-4 border-neutral-300 dark:border-neutral-600 pl-4 italic text-neutral-600 dark:text-neutral-400 my-6"&gt;
&lt;p&gt;The figures on deaths relate in almost all cases to patients who have died in hospital and who have tested positive for COVID-19.[&amp;hellip;] These figures do not include deaths outside hospital, such as those in care homes, except as indicated above.&lt;sup id="fnref:5"&gt;&lt;a href="#fn:5" class="footnote-ref" role="doc-noteref"&gt;5&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;So, again: real figures are much higher than those reported, no matter which country made the measurements. All of them are wrong. Some people argue that governments do not want to provide real figures not to create even more social alarm, lose popularity with their voters or even to look better than other countries. However, often the simpler answer is the most probable one&lt;sup id="fnref:6"&gt;&lt;a href="#fn:6" class="footnote-ref" role="doc-noteref"&gt;6&lt;/a&gt;&lt;/sup&gt;: it is not that governments want to hide information from us, it is just that no country has the means to face this outbreak, nor to mention to take accurate metrics. &lt;strong&gt;And here lies another drama of COVID-19 that goes beyond the personal tragedy: no country in the world is prepared for the stress test that COVID-19 represents.&lt;/strong&gt;&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/conspiracy-mulder.gif"
alt="It is not that Governments are conspiring hiding information, it is just that they do not have the means to get better figures. And here lies the tragedy!"&gt;&lt;figcaption&gt;
&lt;p&gt;It is not that Governments are conspiring hiding information, it is just that they do not have the means to get better figures. And here lies the tragedy!&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;But let&amp;rsquo;s go back to our path: that of data and figures. At this point, we have to acknowledge that all the data is flawed and we cannot get a perfect picture of the real situation out of them. If we wanted to do so, we would need to use other data sources, like comparing the record of total daily deaths&lt;sup id="fnref:7"&gt;&lt;a href="#fn:7" class="footnote-ref" role="doc-noteref"&gt;7&lt;/a&gt;&lt;/sup&gt; with the same period last year(s). Of course, this will need more time, and, in turn, this also has other implications and problems (for example, it will not give an accurate number of deaths by COVID-19, as there is no way to know their cause of death, but the significant difference between periods could be a good proxy).&lt;/p&gt;
&lt;p&gt;So we have two options now, either losing faith completely in all COVID-19 infographics and metrics or to acknowledge their limitations and assume that they are just a rough approximation to reality.&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/morpheus-red-pill.jpg"
alt="We have two options: either losing faith completely in all COVID-19 infographics or assuming they are no more than a rough approximation to reality"&gt;&lt;figcaption&gt;
&lt;p&gt;We have two options: either losing faith completely in all COVID-19 infographics or assuming they are no more than a rough approximation to reality&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="3-choose-the-right-figures-and-visuals"&gt;3. Choose the right figures and visuals&lt;/h2&gt;
&lt;p&gt;Great! If you are reading this it means that you are ok assuming that reality is (as always) far more complex than what nice dashboards can ever show, no matter how fancy they are. And speaking of that: &lt;strong&gt;beware of fancy visuals!&lt;/strong&gt;&lt;/p&gt;
&lt;div class="gallery" style="display: flow-root"&gt;
&lt;a data-fancybox="gallery-img/showcase" href="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/FT-rollingdata2.png" &gt;
&lt;div style="background-image:url(/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/FT-rollingdata2_hu_1f1616995ea14377.png); background-size: cover; background-position: 50%; width: 30%; min-height: 200px; height: auto; float:left; margin: 5px"&gt;
&lt;/div&gt;
&lt;/a&gt;
&lt;a data-fancybox="gallery-img/showcase" href="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/NYT%20Coronavirus%20in%20the%20U%20S%20Latest%20Map%20and%20Case%20Count.png" data-caption="Map of OVID-19 Cases&amp;amp;rsquo; growth in USA. Source:"&gt;
&lt;div style="background-image:url(/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/NYT%20Coronavirus%20in%20the%20U%20S%20Latest%20Map%20and%20Case%20Count_hu_5a5fd1138b4495f5.png); background-size: cover; background-position: 50%; width: 30%; min-height: 200px; height: auto; float:left; margin: 5px"&gt;
&lt;/div&gt;
&lt;/a&gt;
&lt;a data-fancybox="gallery-img/showcase" href="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/Screenshot_2020-04-03%20How%20the%20Virus%20Got%20Out.png" &gt;
&lt;div style="background-image:url(/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/Screenshot_2020-04-03%20How%20the%20Virus%20Got%20Out_hu_5ce9a6ca027a2d3e.png); background-size: cover; background-position: 50%; width: 30%; min-height: 200px; height: auto; float:left; margin: 5px"&gt;
&lt;/div&gt;
&lt;/a&gt;
&lt;a data-fancybox="gallery-img/showcase" href="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/Table-decoration.png" data-caption="A table combining data with (rough) visualization. Source:"&gt;
&lt;div style="background-image:url(/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/Table-decoration_hu_bad38023cded6fd6.png); background-size: cover; background-position: 50%; width: 30%; min-height: 200px; height: auto; float:left; margin: 5px"&gt;
&lt;/div&gt;
&lt;/a&gt;
&lt;a data-fancybox="gallery-img/showcase" href="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/UKs%20COVID-19%20Dashboard.png" data-caption="Total UK COVID-19 Cases Update Source:"&gt;
&lt;div style="background-image:url(/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/UKs%20COVID-19%20Dashboard_hu_2fd06161bcda2cb4.png); background-size: cover; background-position: 50%; width: 30%; min-height: 200px; height: auto; float:left; margin: 5px"&gt;
&lt;/div&gt;
&lt;/a&gt;
&lt;a data-fancybox="gallery-img/showcase" href="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/WHO%20Health%20Emergency%20Dashboard.png" data-caption="WHO map. Source:"&gt;
&lt;div style="background-image:url(/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/WHO%20Health%20Emergency%20Dashboard_hu_bb8a533c588ccc49.png); background-size: cover; background-position: 50%; width: 30%; min-height: 200px; height: auto; float:left; margin: 5px"&gt;
&lt;/div&gt;
&lt;/a&gt;
&lt;a data-fancybox="gallery-img/showcase" href="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/covid19_fallecimientos-por-region-superpuesto-offset-log_since-5deceased.png" &gt;
&lt;div style="background-image:url(/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/covid19_fallecimientos-por-region-superpuesto-offset-log_since-5deceased_hu_d267f7d3f28bc5df.png); background-size: cover; background-position: 50%; width: 30%; min-height: 200px; height: auto; float:left; margin: 5px"&gt;
&lt;/div&gt;
&lt;/a&gt;
&lt;a data-fancybox="gallery-img/showcase" href="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/numeroteca-small-multiples.png" data-caption="Numeroteca&amp;amp;rsquo;s Evolution of cases in Spain&amp;amp;rsquo;s regions, as part of an exhaustive analysis on Spain, France and Italy. Source:"&gt;
&lt;div style="background-image:url(/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/numeroteca-small-multiples_hu_28f23740e063950f.png); background-size: cover; background-position: 50%; width: 30%; min-height: 200px; height: auto; float:left; margin: 5px"&gt;
&lt;/div&gt;
&lt;/a&gt;
&lt;a data-fancybox="gallery-img/showcase" href="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/numeroteca_covid19_casos-registrados-por-comunidad-autonoma-superpuesto-log.png" data-caption="Numeroteca&amp;amp;rsquo;s Evolution of cases in Spain&amp;amp;rsquo;s regions, as part of an exhaustive analysis on Spain, France and Italy. Source:"&gt;
&lt;div style="background-image:url(/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/numeroteca_covid19_casos-registrados-por-comunidad-autonoma-superpuesto-log_hu_222f72e4c88c4194.png); background-size: cover; background-position: 50%; width: 30%; min-height: 200px; height: auto; float:left; margin: 5px"&gt;
&lt;/div&gt;
&lt;/a&gt;
&lt;/div&gt;
&lt;p&gt;I&amp;rsquo;m sure that at this point you may have seen plenty of neat infographics of any type, like the ones above: some of them use boxplots, other lines, other scatterplots&amp;hellip; some of them have smooth edges, others have axis with logarithmic scales. Even there are some that want to introduce geospatial analysis and present maps of different kinds (choropleths, bubbles, sizes)&amp;hellip; and if you are like me, you can enjoy watching them and interacting with them during hours. But are they really effective to display useful data? Unfortunately, most of them are not (even some of my own).&lt;/p&gt;
&lt;p&gt;One of the most basic yet frequent is to display a big figure of the total cases within a country. &lt;strong&gt;Big figures are really catchy and easy to understand, but they usually lack some context to make them really meaningful&lt;/strong&gt;. Of course, anyone can understand that a 7-figure number is a big one, but it is really difficult to know how big it is. We need to compare it to something else to wholly grasp its real magnitude. Also, since we are dealing with a cumulative figure, knowing the analysed time span is a must, as it is not the same to reach a certain figure in one day, one week, one month or one year.&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/big-figures.png"&gt;
&lt;/figure&gt;
&lt;p&gt;A common variation of that is to provide a line plot of cumulative cases, with the dates on the X-axis and total cases in Y-Axis. There are several variations of that: displaying relative data (eg: number of cases/population), with logarithmic scale (in order to make it easier to see the variations of the first days when compared to most recent ones), displaying several categories, either representing different regions or type of cases&amp;hellip; Whereas they are really effective and most of them are right from a technical standpoint (especially considering that some of these variations make a great difference), it is the representation of figure itself that may be of little or no use. Is it really representative of something? What kind of questions can we answer by providing a number that, by definition, will always grow?&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/WHO-cumulative.png"
alt="Line plot displaying cumulative cases. Is it really meaningful? Source:"&gt;&lt;figcaption&gt;
&lt;p&gt;Line plot displaying cumulative cases. Is it really meaningful? Source:&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;It is for that reason that some have decided to use other indicators in order to assess the evolution of the pandemic, such as the number of new cases per day, as they allow us to easily identify if figures are better or worse than the previous day. Again, the same variations of the previous plots can be made in order to make this even more insightful, such as the following barplot, which displays the daily variation of cases, grouped by types. Not only we can see that they are starting to lower, but also, that the number of recovered cases is increasing over the deaths or active cases.&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/cases_day_type.png"
alt="Daily cases, by type. Source:"&gt;&lt;figcaption&gt;
&lt;p&gt;Daily cases, by type. Source:&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This is indeed more useful than cumulative cases, although it is somewhat volatile and can lead to confusion. Take the image below, as an example:&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/weekend-effect2.png"
alt="Can you see that weird unexpected variation? It is the weekend effect! (That&amp;rsquo;s why in the next dashboard version I will be highlighting the weekends). Source:"&gt;&lt;figcaption&gt;
&lt;p&gt;Can you see that weird unexpected variation? It is the weekend effect! (That&amp;rsquo;s why in the next dashboard version I will be highlighting the weekends). Source:&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Can you see how all figures suffer from a dramatic fall on 29th and 30th March just before peaking again the day after? This is indeed an unexpected behaviour and really difficult to explain. Unless we realise that those days were Saturday and Sunday, and due to the fact that there are fewer people working at hospitals, data is not taken as fast as usual and therefore, accumulates on Monday. This phenomenon has been called &amp;ldquo;the weekend effect&amp;rdquo; (Did I mention that you should not take data too seriously? &amp;#x1f609;)&lt;/p&gt;
&lt;p&gt;It is because of that that some others, such as
from Financial Times, prefer to use a rolling average of a fixed period (such as 3 or 5 days), which is a more stable figure, such as that shown in the figure below.&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/FT-rollingdata.png"
alt="Source:"&gt;&lt;figcaption&gt;
&lt;p&gt;Source:&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Other techniques used to fix those outliers and display tendencies are to use smooth line plots based on the actual data, like the following plot made by
.&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/numeroteca-smooth.jpeg"
alt="Smoothed lines based on actual data. Source:"&gt;&lt;figcaption&gt;
&lt;p&gt;Smoothed lines based on actual data. Source:&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="4-do-not-make-hasty-comparisons"&gt;4. Do not make hasty comparisons&lt;/h2&gt;
&lt;p&gt;Surely, the most frequent type of visual is that comparing how COVID-19 is affecting different regions, either within a country or comparing different countries (usually using China or Wuchan as a reference -after all, it is where it all started). While this kind of plots could provide answers to questions such as how a specific region is doing regarding another one (and therefore, replicating or avoiding their measures against COVID-19, for example), those comparisons are really problematic. For starters, the fact that population or size is largely different invalidates any comparison in absolute terms.&lt;/p&gt;
&lt;p&gt;But even when using relative values (eg: number of cases per inhabitant), there are other key factors that have a direct impact on the evolution of the disease and its effects are assumed to be the same, while in reality can differ in several orders of magnitude, such as demography&lt;sup id="fnref:8"&gt;&lt;a href="#fn:8" class="footnote-ref" role="doc-noteref"&gt;8&lt;/a&gt;&lt;/sup&gt;, geography, urban settlements or health systems (in terms of human and financial resources), just to name a few.&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/Diario_es_demografia_infectados2.png"
alt="Histogram of cases and mortality rates by age and country. Most cases in Italy were from 70-90, as opposed to 20-30 in South Korea, hence the enormous difference in the total of deaths. Source:"&gt;&lt;figcaption&gt;
&lt;p&gt;Histogram of cases and mortality rates by age and country. Most cases in Italy were from 70-90, as opposed to 20-30 in South Korea, hence the enormous difference in the total of deaths. Source:&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;strong&gt;Amongst all those differences, though, the most relevant well may be the number of tests done to detect COVID-19, a figure that is not just different but also usually unknown&lt;/strong&gt;. This is by no means trivial, as we have seen that the number of cases is defined using this parameter. Therefore, if a country performed a really low number of tests, it will also have an extremely low number of cases. Out of sight, out of mind.&lt;/p&gt;
&lt;p&gt;So, the biggest problem here is that without taking into consideration those factors, comparisons may render plenty of biased conclusions that have nothing to do with reality, such as come countries may be dodging COVID-19 or are kind of immune, or that COVID-19 only affects countries with a bad health system, unorganized governments, or undeveloped countries. As a result, &lt;strong&gt;there is a risk of developing a narrative of moral superiority&lt;sup id="fnref:9"&gt;&lt;a href="#fn:9" class="footnote-ref" role="doc-noteref"&gt;9&lt;/a&gt;&lt;/sup&gt; based on totally wrong foundations&lt;/strong&gt; like what some politicians have started to do in their own self-interest&lt;sup id="fnref:10"&gt;&lt;a href="#fn:10" class="footnote-ref" role="doc-noteref"&gt;10&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/repugnant-conclusions.jpg"
alt="Don&amp;rsquo;t be like Wopke Hoekstra doing hasty comparison, or you risk reaching repugnant conclusions like him. Photo:"&gt;&lt;figcaption&gt;
&lt;p&gt;Don&amp;rsquo;t be like Wopke Hoekstra doing hasty comparison, or you risk reaching repugnant conclusions like him. Photo:&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="5-take-predictions-even-less-seriously"&gt;5. Take predictions even less seriously&lt;/h2&gt;
&lt;p&gt;The last group of plots, and the most complex ones, are those that make predictions. While they are really appealing and, apparently, provide answers to one of our main concerns (&lt;em&gt;&amp;ldquo;When is this going to end?&amp;rdquo; / Will this last any longer?&lt;/em&gt;) in a very understandable way, they are really tricky. There are several ways to make predictions, such as using linear regression or models. While a model can be as easy&lt;sup id="fnref:11"&gt;&lt;a href="#fn:11" class="footnote-ref" role="doc-noteref"&gt;11&lt;/a&gt;&lt;/sup&gt; or as complex as we want it to be (and as a result, their accuracy will differ dramatically), they mostly rely on having a good set of historic data or knowing the logics of the phenomenon they want to describe. Unfortunately, since COVID-19 is a new phenomenon, we are lacking of both, and therefore, predictions at this stage are prone to errors. Some predictions are based on what has happened in other places where the outbreak started before, but we have seen how problematic this can be.&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/tarot.jpeg"
alt="If we can&amp;rsquo;t see (nor understand) the model behind a prediction, it can be as useful and reliable as that of a fortune teller."&gt;&lt;figcaption&gt;
&lt;p&gt;If we can&amp;rsquo;t see (nor understand) the model behind a prediction, it can be as useful and reliable as that of a fortune teller.&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Last, but not least, we should not forget that they require a large amount of knowledge of a particular field something like most of the people who are doing those nice visualizations (including myself) lack of. Therefore, since I have openly admitted that this is something beyond my knowledge, at this point I can only but recommend to be sceptical about any prediction that is not made by an authority on the field. And even in that case (provided I could understand it), I would recommend caution. Or even better: just rely on predictions if they were made by epidemiologists and you are one of them.&lt;/p&gt;
&lt;h2 id="wrapping-up"&gt;Wrapping up&lt;/h2&gt;
&lt;p&gt;As argued, if data visualization is never easy, it is even less so in the case of a novel phenomenon such as COVID-19. Therefore, when facing any type of visuals, we should proceed with caution. If you are doing (or planning to do) any type of visualization, ask yourself what question do you want to answer and which is the best way to do it, take into account the aforementioned considerations and make them evident to your readers&lt;sup id="fnref:12"&gt;&lt;a href="#fn:12" class="footnote-ref" role="doc-noteref"&gt;12&lt;/a&gt;&lt;/sup&gt;. Also, make your analysis reproducible, so anyone could tell you if you did something wrong or even fix it by themselves. If you are simply watching them, look for all those explanations, and if you can&amp;rsquo;t find them, ask for them, help the author or simply ignore it and look for an alternative. But in any case, you should always remember not to take data too seriously or too blindly. Data by itself is not what really matters, is what we do with it and how we do it in order to achieve knowledge what really matters. And here&amp;rsquo;s when science plays its role.&lt;/p&gt;
&lt;div class="footnotes" role="doc-endnotes"&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id="fn:1"&gt;
&lt;p&gt;There is a third scenario: to infer or calculate the data from other datasets.&amp;#160;&lt;a href="#fnref:1" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:2"&gt;
&lt;p&gt;For example,
, without a clean structure. Not only they use a proprietary format, but they mix data with metadata on the same file, even in the same sheet, instead of using a wide format where every column is a field and every row is an observation or a long format where keys and values are stored as in a dictionary. On the other hand, Spain decided to release the data in PDF. PDFs are great for visualization because they are an ISO format and can be opened with plenty of different softwares. Unfortunately, it is not great for consuming data, as it is not structured in any way.&amp;#160;&lt;a href="#fnref:2" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:3"&gt;
&lt;p&gt;At the time of writing this post, and according to
&amp;rsquo;s page, the
. The standard method of testing is real-time reverse transcription polymerase chain reaction (rRT-PCR), typically done on respiratory samples obtained by a nasopharyngeal swab and results are generally available within a few hours to two days.&amp;#160;&lt;a href="#fnref:3" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:4"&gt;
&lt;p&gt;This article from El Pais (in Spanish)
&amp;#160;&lt;a href="#fnref:4" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:5"&gt;
&lt;p&gt;
&amp;#160;&lt;a href="#fnref:5" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:6"&gt;
&lt;p&gt;This is the commonly phrased version of
principle.&amp;#160;&lt;a href="#fnref:6" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:7"&gt;
&lt;p&gt;In Spain
&amp;#160;&lt;a href="#fnref:7" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:8"&gt;
&lt;p&gt;Refer to this (
-in Spanish) which focuses on the fact that coronavirus mortality rate differs enormously according to the age.&amp;#160;&lt;a href="#fnref:8" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:9"&gt;
&lt;p&gt;Refer to
(Publico, in Spanish)&amp;#160;&lt;a href="#fnref:9" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:10"&gt;
&lt;p&gt;Take as an example the unfortunate words by Dutch finance minister Wopke Hoekstra, who suggested the EU &amp;ldquo;should investigate countries like Spain that say they have no budgetary margin to deal with the effects of the crisis provoked by the new coronavirus in spite of the fact that the eurozone has grown for seven consecutive years&amp;rdquo;, a statement that was later qualified as &amp;ldquo;repugnant&amp;rdquo; by Portugal&amp;rsquo;s Primer Minister, António Costa. (Source:
)&amp;#160;&lt;a href="#fnref:10" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:11"&gt;
&lt;p&gt;Just to give you an example of how easy can be to implement a prediction in R, refer to this article:
&amp;#160;&lt;a href="#fnref:11" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:12"&gt;
&lt;p&gt;Also, reading this post may be useful:
&amp;#160;&lt;a href="#fnref:12" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</description></item></channel></rss>