<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>R |</title><link>https://carlos-hugoblox.netlify.app/en/tags/r/</link><atom:link href="https://carlos-hugoblox.netlify.app/en/tags/r/index.xml" rel="self" type="application/rss+xml"/><description>R</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-GB</language><lastBuildDate>Thu, 01 Sep 2022 13:00:00 +0000</lastBuildDate><image><url>https://carlos-hugoblox.netlify.app/media/icon_hu_aa3341e371185529.png</url><title>R</title><link>https://carlos-hugoblox.netlify.app/en/tags/r/</link></image><item><title>Quarto: a library to run them all? A collaborative exercise to use, learn and assess quarto for authoring reproducible documents in different scenarios</title><link>https://carlos-hugoblox.netlify.app/en/events/2022-09-07-rseconf22/</link><pubDate>Thu, 01 Sep 2022 13:00:00 +0000</pubDate><guid>https://carlos-hugoblox.netlify.app/en/events/2022-09-07-rseconf22/</guid><description>&lt;!-- Add the talk outline, prerequisites, and how people can join. --&gt;
&lt;p&gt;Using literate programming is a widespread practice amongst data scientists. This practice not only encourages data scientists to produce transparent, rich and reflective accounts of their analysis without the extra overhead of switching between tools, but also leads to artefacts (i.e., notebooks) that are increasingly becoming a medium for dissemination, reproducibility and education. Rmarkdown or Jupyter notebooks, are two of the most well known and used options. While both solutions can be used with multiple programming languages, the decision of whether to use one or the other is almost certain to be exclusively based on that. At least until now, with Quarto being mature enough to become a game-changer.&lt;/p&gt;
&lt;p&gt;Quarto is a language-agnostic software based on Pandoc to render files combining markdown and code into multiple ranges of formats and outputs. As a result, it can be used with either R, Python or Julia without any other dependencies.&lt;/p&gt;
&lt;p&gt;This collaborative workshop will be structured as follows:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;a brief theoretical introduction and instructions;&lt;/li&gt;
&lt;li&gt;a task that participants may choose from the different use cases provided (i.e. generating a single document, migrating from Rmarkdown or Jupyter, creating a book or generating interactive content); and&lt;/li&gt;
&lt;li&gt;a group discussion and conclusions.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Participants in this workshop have fun while gaining enough depth of breath and practice to evaluate how feasible it is to use Quarto in different scenarios and, ultimately, if it can become the one tool for authoring reproducible scientific or technical documents, regardless of your language of choice.&lt;/p&gt;
&lt;h2 id="expertise-level"&gt;Expertise level&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Practitioner&lt;/li&gt;
&lt;li&gt;Expert&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="audience"&gt;Audience&lt;/h2&gt;
&lt;p&gt;This workshop is targeted to anyone interested in reproducible publishing, especially if they are looking for a single solution to use regardless of their programming language and are willing to explore with others quarto&amp;rsquo;s pros and cons, contributing to a group discussion.&lt;/p&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;
&lt;img alt="Photo by Tyneshight"
srcset="https://carlos-hugoblox.netlify.app/en/events/2022-09-07-rseconf22/RSEConf-070922-79_hu_c5885ac3bf037022.webp 320w, https://carlos-hugoblox.netlify.app/en/events/2022-09-07-rseconf22/RSEConf-070922-79_hu_6286f0458a951e69.webp 480w, https://carlos-hugoblox.netlify.app/en/events/2022-09-07-rseconf22/RSEConf-070922-79_hu_2fb86d591a05df0c.webp 760w"
sizes="(max-width: 480px) 100vw, (max-width: 768px) 90vw, (max-width: 1024px) 80vw, 760px"
src="https://carlos-hugoblox.netlify.app/en/events/2022-09-07-rseconf22/RSEConf-070922-79_hu_c5885ac3bf037022.webp"
width="760"
height="506"
loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;</description></item><item><title>Managing R script dependencies: automagic and renv</title><link>https://carlos-hugoblox.netlify.app/en/events/2022-07-14-wrug-reproducbility/</link><pubDate>Thu, 14 Jul 2022 13:00:06 +0000</pubDate><guid>https://carlos-hugoblox.netlify.app/en/events/2022-07-14-wrug-reproducbility/</guid><description>&lt;!-- Add the talk outline, prerequisites, and how people can join. --&gt;
&lt;p&gt;At the end of the
we had an interesting discussion on the need for reproducibility. This talk introduces the concept of reproducibility and focuses on one of it’s many facets: that of managing R script’s dependencies.&lt;/p&gt;
&lt;p&gt;In order to do so, two different methods will be presented, compared and discussed: the simple yet works-out-of-the-box {automagic} and the more complex, yet backed up by RStudio, {renv}&lt;/p&gt;
&lt;p&gt;At the end of the session, points in common as well as limitations will be highlighted, hopefully leading to a discussion and opening the door for the second talk of the reproducibility series.&lt;/p&gt;</description></item><item><title>Visualisation tool for a grounded FEW Nexus</title><link>https://carlos-hugoblox.netlify.app/en/publications/2022-fwe-dashboard/</link><pubDate>Fri, 20 May 2022 00:00:00 +0000</pubDate><guid>https://carlos-hugoblox.netlify.app/en/publications/2022-fwe-dashboard/</guid><description>&lt;!-- Add the paper text or supplementary notes. Markdown, math, and code are supported. --&gt;</description></item><item><title>Streetscape Perception Modelling – Theoretical Considerations and Methodological Possibilities</title><link>https://carlos-hugoblox.netlify.app/en/events/2021-12-15-platial/</link><pubDate>Fri, 17 Dec 2021 13:00:06 +0000</pubDate><guid>https://carlos-hugoblox.netlify.app/en/events/2021-12-15-platial/</guid><description>&lt;!-- Add the talk outline, prerequisites, and how people can join. --&gt;
&lt;p&gt;The influence of features and properties of the urban built-up environment on people’s sense of safety and perception of beauty, social vibrancy, and walkability is a topic of interest of urban geographers, designers, planners, and environmental psychologists alike. Along with emerging forms of data and the computational paradigm of Artificial Intelligence, current GIS, citizen science, and sensor technologies offer exciting technical and methodological possibilities for extracting, representing, and modelling aspects of people’s perception of streetscapes. This workshop aims to explore these possibilities, exchange research experiences, and discuss the theoretical grounds based on which we can operationalise, i.e., model, streetscape perception, particularly based on geospatial technologies.&lt;/p&gt;
&lt;p&gt;The workshop will offer an open and interactive environment for researchers of all levels to discuss questions including, but not limited to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;What theoretical frameworks can guide us in extracting, representing, and modelling different streetscape perception aspects of people from different social groups and cultural backgrounds?&lt;/li&gt;
&lt;li&gt;Which variables, parameters, and indicators these theoretical frameworks suggest and how can they be reliably extracted based on different digital technologies and participation methods?&lt;/li&gt;
&lt;li&gt;Which GIS interfaces, data structures, and visualisations are effective in representing people’s perception of streetscapes as to foster theory development and inform planning and policy design towards more sustainable and inclusive urban spaces?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;An outcome of the workshop will be a written summary of the discussed ideas authored by all participants and published in the PLATIAL’21 proceedings. The indicative programme is as follows:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Introduction by the organisers (20 min)&lt;/li&gt;
&lt;li&gt;Crash presentations by participants (optional) (20 min)&lt;/li&gt;
&lt;li&gt;Break-out group discussions (30 min) + World Café (30 min)&lt;/li&gt;
&lt;li&gt;Paper writing outlook (20 min)&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title>Creating Interfaces</title><link>https://carlos-hugoblox.netlify.app/en/projects/creating-interfaces/</link><pubDate>Mon, 11 Oct 2021 00:00:00 +0000</pubDate><guid>https://carlos-hugoblox.netlify.app/en/projects/creating-interfaces/</guid><description>&lt;p&gt;An open source prototype for a visual interface to support research and Nexus engagements, designed collaborativelly as part of
WP4, developed by the
at the
.&lt;/p&gt;
&lt;h2 id="aim"&gt;Aim&lt;/h2&gt;
&lt;p&gt;The aim of this tool is to provide an interface capable of understanding the implications of our decisions regarding food, and how can meals in kindergartens be turned into drivers for positive change for the health and the environment.&lt;/p&gt;
&lt;p&gt;The visualisation tool may be helpful to perform tasks such as:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Identify highly rated meals and the lower rated meals&lt;/li&gt;
&lt;li&gt;Identify meals with higher footprint&lt;/li&gt;
&lt;li&gt;Identify ingredients with higher footprint&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Which, ultimately, will lead to discussions and reflections on how food, energy and water are interlinked and how small changes in food can make a big impact.&lt;/p&gt;
&lt;h2 id="online-demos"&gt;Online demos&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Food in Kindergartens, Slupsk (Poland):
&lt;/li&gt;
&lt;li&gt;Food waste, Wilmington (USA):
&lt;/li&gt;
&lt;li&gt;Local producers, Tulcea (Romania):
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="source-code"&gt;Source Code&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Source Code:
&lt;/li&gt;
&lt;li&gt;Licence: GPL&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="acknowledgements"&gt;Acknowledgements&lt;/h2&gt;
&lt;p&gt;This study was conducted in the Creating Interfaces project, funded within the framework of the Sustainable Global Urban Initiative (SUGI) Food-Water-Energy Nexus program. This program has been set up by the Belmont Forum and the Joint Programming Initiative (JPI) Urban Europe and has received funding from the European Union&amp;rsquo;s Horizon, 2020 research and innovation program under grant agreement # 730254 and the following national funding agencies: United Kingdom Research and Innovation funding was received through the Economic and Social Science Research Council (ESRC) grant ES/S002235/1; the National Science Center (NCN) of Poland funded this work under grant #UMO-2017/25/Z/HS6/03046.&lt;/p&gt;</description></item><item><title>Considerations on the importance of data and science in data science</title><link>https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/</link><pubDate>Sun, 05 Apr 2020 00:00:00 +0000</pubDate><guid>https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/</guid><description>&lt;!-- Tip: open with the why, then show results, code, and next steps. --&gt;
&lt;p&gt;I must confess: during these days of lockdown, I have been toying with dashboards and infographics about COVID-19 outbreak, just like everyone else.
with nice plots, some formatted tables and even some (basic) maps. It&amp;rsquo;s shiny and nice, you can spend some time playing with it turning layers on and off, hovering, zooming&amp;hellip; And, honestly, I have to admit that I am proud of some of the results I have achieved. &lt;strong&gt;And yet, it is flawed. Just like everyone else&amp;rsquo;s&lt;/strong&gt; (or almost). And yet, I will keep improving it, even though I&amp;rsquo;m afraid it will always be flawed and even though I acknowledge that it will never be a contribution to improve knowledge on the topic.&lt;/p&gt;
&lt;p&gt;Why, then, am I persisting on keeping working on it if I know I cannot change its fate? Admittedly, at some point, I asked myself that very question and I even considered quitting. Not only I didn&amp;rsquo;t want to lose my time (even in these days when we are locked down at home there are plenty of things we can do), but I didn&amp;rsquo;t want to contribute to generating noise, misinformation and even more dramatism about an already important drama. Because that&amp;rsquo;s what flawed graphics do. But in the end, I realised that &lt;strong&gt;working on a dashboard like that could be a great opportunity for learning by doing.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Many types of lessons can be learnt from it, like those related to technical skills, or those related to how data is gathered, visualized and analyzed. &lt;strong&gt;Today, when data and figures on COVID-19 are everywhere, I want to share some reflections on the science (or lack of it) in data science.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="1-get-the-right-data"&gt;1. Get the (right) data.&lt;/h2&gt;
&lt;p&gt;As obvious as it sounds, there is no data visualization nor data science without data to visualize or analyze. Therefore, the first thing that anyone who wants to make any type of data visualization or data analysis is to get the data. Second: we cannot use any type of data. We need to use &lt;em&gt;good data&lt;/em&gt;, and &lt;strong&gt;by &lt;em&gt;good&lt;/em&gt; I mean &lt;em&gt;reliable&lt;/em&gt;, &lt;em&gt;usable&lt;/em&gt; (in terms of licences, formats and structure), up-to-date and frequently &lt;em&gt;updated&lt;/em&gt;, and, hopefully, &lt;em&gt;official&lt;/em&gt; data that is &lt;em&gt;representative&lt;/em&gt; enough to explain the phenomenon we are studying.&lt;/strong&gt; While this is usually non-trivial, it is even more crucial if we are to explain a completely new phenomenon that it is happening as we speak and it does at a global scale like COVID-19.&lt;/p&gt;
&lt;p&gt;Usually, there are only two possible options&lt;sup id="fnref:1"&gt;&lt;a href="#fn:1" class="footnote-ref" role="doc-noteref"&gt;1&lt;/a&gt;&lt;/sup&gt;: either gather data by ourselves or rely on others&amp;rsquo; data. Whereas gathering our own data might be the best choice for some scenarios, in the case of COVID-19, it is unlikely that we are in a position to gather the kind of data that might be useful for us (in fact, even governments are struggling to do so, as we will see). Therefore, we are left to just one option. Of course, we cannot rely on some random person or institution, we need to rely on someone we can trust, like universities (because they tend to provide rigorous data), governments (because they provide official data) or organizations (like the
). But where do we get the data from?&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/WHO-dashboard.png"
alt="WHO&amp;rsquo;s COVID-19 dashboard (Screenshot from 02-04-2020)"&gt;&lt;figcaption&gt;
&lt;p&gt;WHO&amp;rsquo;s COVID-19 dashboard (Screenshot from 02-04-2020)&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;While most governments provide data in open licences that allow its use and reuse for any purpose, they usually fail to address another important issue: the file format and data structure. As a result, in countries like the UK or Spain, data is not ready to consume as it is, without manual work&lt;sup id="fnref:2"&gt;&lt;a href="#fn:2" class="footnote-ref" role="doc-noteref"&gt;2&lt;/a&gt;&lt;/sup&gt;. Others, such as WHO, provide nice dashboards with data, but they do not provide the raw data to be consumed by 3rd parties. The good news is that there are people and institutions who are working in releasing clean data with open licences, such as
,
that gathers data from different governments and WHO, and even
that fetches data from UJH and provides a nice data frame for developers and data scientists to use in their projects.&lt;/p&gt;
&lt;p&gt;As a result, it is no wonder that most infographics and dashboards worldwide rely on the same data sources. It seems a sensible decision: not only we get data which is ready to use, but we do it from trustful sources. And yet, as I will argue, it is because of that reason that most of them are wrong. But what could possibly go wrong?&lt;/p&gt;
&lt;h2 id="2-dont-take-data-too-seriously"&gt;2. Don&amp;rsquo;t take data too seriously&lt;/h2&gt;
&lt;p&gt;Now that we know where to get the data from, there is something we have to be aware of: by relying on data generated by others we have not solved the (main) problem of data gathering, we have simply transferred the responsibility to somebody else, but someone still has to deal with what we have been trying to avoid. And, surprise, not every country gathers the data in the same way.&lt;/p&gt;
&lt;p&gt;Take the case of the most basic and crucial question: &lt;strong&gt;how are the number of confirmed cases defined&lt;/strong&gt;. Since COVID-19&amp;rsquo;s symptoms are very similar to those of influenza and the only way to know if someone is infected by it is by testing positive in the tests&lt;sup id="fnref:3"&gt;&lt;a href="#fn:3" class="footnote-ref" role="doc-noteref"&gt;3&lt;/a&gt;&lt;/sup&gt;. This seems a great definition: we have an objective test which is same the for everyone and all countries seem to use the same criteria. Unfortunately, those tests require equipment which is scarce (compared to the current worldwide demand), can only be made in hospitals, and require up to two days to get the results. Therefore, there are many other scenarios that are not considered within this test, such as those shown in figure 2a. &lt;strong&gt;So yes, every government provides that figure, yet all of them are much lower than the real figure. How much lower? There is no way to know.&lt;/strong&gt;&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/iceberg.png"
alt="Confirmed figures ≠ actual figures. There are many more cases than the official ones. How many more? We can&amp;rsquo;t possibly know"&gt;&lt;figcaption&gt;
&lt;p&gt;Confirmed figures ≠ actual figures. There are many more cases than the official ones. How many more? We can&amp;rsquo;t possibly know&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Let&amp;rsquo;s focus on another example: &lt;strong&gt;the number of deaths by COVID 19&lt;/strong&gt;. Apparently, this should be easier. Every country keeps a record on the number of deaths per day and its cause of death, so it should be easy to filter those who died from COVID-19 amongst all the possible causes. Well, no. As we have seen, if we can&amp;rsquo;t define with precision the number of people who are infected by COVID-19, we will not be able to know the number of people who have died as a result of it.&lt;/p&gt;
&lt;p&gt;But it can be even trickier if we consider that every country has different criteria on how they count the number of deaths of people who were tested positive&lt;sup id="fnref:4"&gt;&lt;a href="#fn:4" class="footnote-ref" role="doc-noteref"&gt;4&lt;/a&gt;&lt;/sup&gt;. Take the case of UK&amp;rsquo;s definition:&lt;/p&gt;
&lt;blockquote class="border-l-4 border-neutral-300 dark:border-neutral-600 pl-4 italic text-neutral-600 dark:text-neutral-400 my-6"&gt;
&lt;p&gt;The figures on deaths relate in almost all cases to patients who have died in hospital and who have tested positive for COVID-19.[&amp;hellip;] These figures do not include deaths outside hospital, such as those in care homes, except as indicated above.&lt;sup id="fnref:5"&gt;&lt;a href="#fn:5" class="footnote-ref" role="doc-noteref"&gt;5&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;So, again: real figures are much higher than those reported, no matter which country made the measurements. All of them are wrong. Some people argue that governments do not want to provide real figures not to create even more social alarm, lose popularity with their voters or even to look better than other countries. However, often the simpler answer is the most probable one&lt;sup id="fnref:6"&gt;&lt;a href="#fn:6" class="footnote-ref" role="doc-noteref"&gt;6&lt;/a&gt;&lt;/sup&gt;: it is not that governments want to hide information from us, it is just that no country has the means to face this outbreak, nor to mention to take accurate metrics. &lt;strong&gt;And here lies another drama of COVID-19 that goes beyond the personal tragedy: no country in the world is prepared for the stress test that COVID-19 represents.&lt;/strong&gt;&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/conspiracy-mulder.gif"
alt="It is not that Governments are conspiring hiding information, it is just that they do not have the means to get better figures. And here lies the tragedy!"&gt;&lt;figcaption&gt;
&lt;p&gt;It is not that Governments are conspiring hiding information, it is just that they do not have the means to get better figures. And here lies the tragedy!&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;But let&amp;rsquo;s go back to our path: that of data and figures. At this point, we have to acknowledge that all the data is flawed and we cannot get a perfect picture of the real situation out of them. If we wanted to do so, we would need to use other data sources, like comparing the record of total daily deaths&lt;sup id="fnref:7"&gt;&lt;a href="#fn:7" class="footnote-ref" role="doc-noteref"&gt;7&lt;/a&gt;&lt;/sup&gt; with the same period last year(s). Of course, this will need more time, and, in turn, this also has other implications and problems (for example, it will not give an accurate number of deaths by COVID-19, as there is no way to know their cause of death, but the significant difference between periods could be a good proxy).&lt;/p&gt;
&lt;p&gt;So we have two options now, either losing faith completely in all COVID-19 infographics and metrics or to acknowledge their limitations and assume that they are just a rough approximation to reality.&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/morpheus-red-pill.jpg"
alt="We have two options: either losing faith completely in all COVID-19 infographics or assuming they are no more than a rough approximation to reality"&gt;&lt;figcaption&gt;
&lt;p&gt;We have two options: either losing faith completely in all COVID-19 infographics or assuming they are no more than a rough approximation to reality&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="3-choose-the-right-figures-and-visuals"&gt;3. Choose the right figures and visuals&lt;/h2&gt;
&lt;p&gt;Great! If you are reading this it means that you are ok assuming that reality is (as always) far more complex than what nice dashboards can ever show, no matter how fancy they are. And speaking of that: &lt;strong&gt;beware of fancy visuals!&lt;/strong&gt;&lt;/p&gt;
&lt;div class="gallery" style="display: flow-root"&gt;
&lt;a data-fancybox="gallery-img/showcase" href="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/FT-rollingdata2.png" &gt;
&lt;div style="background-image:url(/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/FT-rollingdata2_hu_1f1616995ea14377.png); background-size: cover; background-position: 50%; width: 30%; min-height: 200px; height: auto; float:left; margin: 5px"&gt;
&lt;/div&gt;
&lt;/a&gt;
&lt;a data-fancybox="gallery-img/showcase" href="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/NYT%20Coronavirus%20in%20the%20U%20S%20Latest%20Map%20and%20Case%20Count.png" data-caption="Map of OVID-19 Cases&amp;amp;rsquo; growth in USA. Source:"&gt;
&lt;div style="background-image:url(/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/NYT%20Coronavirus%20in%20the%20U%20S%20Latest%20Map%20and%20Case%20Count_hu_5a5fd1138b4495f5.png); background-size: cover; background-position: 50%; width: 30%; min-height: 200px; height: auto; float:left; margin: 5px"&gt;
&lt;/div&gt;
&lt;/a&gt;
&lt;a data-fancybox="gallery-img/showcase" href="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/Screenshot_2020-04-03%20How%20the%20Virus%20Got%20Out.png" &gt;
&lt;div style="background-image:url(/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/Screenshot_2020-04-03%20How%20the%20Virus%20Got%20Out_hu_5ce9a6ca027a2d3e.png); background-size: cover; background-position: 50%; width: 30%; min-height: 200px; height: auto; float:left; margin: 5px"&gt;
&lt;/div&gt;
&lt;/a&gt;
&lt;a data-fancybox="gallery-img/showcase" href="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/Table-decoration.png" data-caption="A table combining data with (rough) visualization. Source:"&gt;
&lt;div style="background-image:url(/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/Table-decoration_hu_bad38023cded6fd6.png); background-size: cover; background-position: 50%; width: 30%; min-height: 200px; height: auto; float:left; margin: 5px"&gt;
&lt;/div&gt;
&lt;/a&gt;
&lt;a data-fancybox="gallery-img/showcase" href="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/UKs%20COVID-19%20Dashboard.png" data-caption="Total UK COVID-19 Cases Update Source:"&gt;
&lt;div style="background-image:url(/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/UKs%20COVID-19%20Dashboard_hu_2fd06161bcda2cb4.png); background-size: cover; background-position: 50%; width: 30%; min-height: 200px; height: auto; float:left; margin: 5px"&gt;
&lt;/div&gt;
&lt;/a&gt;
&lt;a data-fancybox="gallery-img/showcase" href="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/WHO%20Health%20Emergency%20Dashboard.png" data-caption="WHO map. Source:"&gt;
&lt;div style="background-image:url(/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/WHO%20Health%20Emergency%20Dashboard_hu_bb8a533c588ccc49.png); background-size: cover; background-position: 50%; width: 30%; min-height: 200px; height: auto; float:left; margin: 5px"&gt;
&lt;/div&gt;
&lt;/a&gt;
&lt;a data-fancybox="gallery-img/showcase" href="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/covid19_fallecimientos-por-region-superpuesto-offset-log_since-5deceased.png" &gt;
&lt;div style="background-image:url(/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/covid19_fallecimientos-por-region-superpuesto-offset-log_since-5deceased_hu_d267f7d3f28bc5df.png); background-size: cover; background-position: 50%; width: 30%; min-height: 200px; height: auto; float:left; margin: 5px"&gt;
&lt;/div&gt;
&lt;/a&gt;
&lt;a data-fancybox="gallery-img/showcase" href="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/numeroteca-small-multiples.png" data-caption="Numeroteca&amp;amp;rsquo;s Evolution of cases in Spain&amp;amp;rsquo;s regions, as part of an exhaustive analysis on Spain, France and Italy. Source:"&gt;
&lt;div style="background-image:url(/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/numeroteca-small-multiples_hu_28f23740e063950f.png); background-size: cover; background-position: 50%; width: 30%; min-height: 200px; height: auto; float:left; margin: 5px"&gt;
&lt;/div&gt;
&lt;/a&gt;
&lt;a data-fancybox="gallery-img/showcase" href="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/numeroteca_covid19_casos-registrados-por-comunidad-autonoma-superpuesto-log.png" data-caption="Numeroteca&amp;amp;rsquo;s Evolution of cases in Spain&amp;amp;rsquo;s regions, as part of an exhaustive analysis on Spain, France and Italy. Source:"&gt;
&lt;div style="background-image:url(/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/showcase/numeroteca_covid19_casos-registrados-por-comunidad-autonoma-superpuesto-log_hu_222f72e4c88c4194.png); background-size: cover; background-position: 50%; width: 30%; min-height: 200px; height: auto; float:left; margin: 5px"&gt;
&lt;/div&gt;
&lt;/a&gt;
&lt;/div&gt;
&lt;p&gt;I&amp;rsquo;m sure that at this point you may have seen plenty of neat infographics of any type, like the ones above: some of them use boxplots, other lines, other scatterplots&amp;hellip; some of them have smooth edges, others have axis with logarithmic scales. Even there are some that want to introduce geospatial analysis and present maps of different kinds (choropleths, bubbles, sizes)&amp;hellip; and if you are like me, you can enjoy watching them and interacting with them during hours. But are they really effective to display useful data? Unfortunately, most of them are not (even some of my own).&lt;/p&gt;
&lt;p&gt;One of the most basic yet frequent is to display a big figure of the total cases within a country. &lt;strong&gt;Big figures are really catchy and easy to understand, but they usually lack some context to make them really meaningful&lt;/strong&gt;. Of course, anyone can understand that a 7-figure number is a big one, but it is really difficult to know how big it is. We need to compare it to something else to wholly grasp its real magnitude. Also, since we are dealing with a cumulative figure, knowing the analysed time span is a must, as it is not the same to reach a certain figure in one day, one week, one month or one year.&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/big-figures.png"&gt;
&lt;/figure&gt;
&lt;p&gt;A common variation of that is to provide a line plot of cumulative cases, with the dates on the X-axis and total cases in Y-Axis. There are several variations of that: displaying relative data (eg: number of cases/population), with logarithmic scale (in order to make it easier to see the variations of the first days when compared to most recent ones), displaying several categories, either representing different regions or type of cases&amp;hellip; Whereas they are really effective and most of them are right from a technical standpoint (especially considering that some of these variations make a great difference), it is the representation of figure itself that may be of little or no use. Is it really representative of something? What kind of questions can we answer by providing a number that, by definition, will always grow?&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/WHO-cumulative.png"
alt="Line plot displaying cumulative cases. Is it really meaningful? Source:"&gt;&lt;figcaption&gt;
&lt;p&gt;Line plot displaying cumulative cases. Is it really meaningful? Source:&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;It is for that reason that some have decided to use other indicators in order to assess the evolution of the pandemic, such as the number of new cases per day, as they allow us to easily identify if figures are better or worse than the previous day. Again, the same variations of the previous plots can be made in order to make this even more insightful, such as the following barplot, which displays the daily variation of cases, grouped by types. Not only we can see that they are starting to lower, but also, that the number of recovered cases is increasing over the deaths or active cases.&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/cases_day_type.png"
alt="Daily cases, by type. Source:"&gt;&lt;figcaption&gt;
&lt;p&gt;Daily cases, by type. Source:&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This is indeed more useful than cumulative cases, although it is somewhat volatile and can lead to confusion. Take the image below, as an example:&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/weekend-effect2.png"
alt="Can you see that weird unexpected variation? It is the weekend effect! (That&amp;rsquo;s why in the next dashboard version I will be highlighting the weekends). Source:"&gt;&lt;figcaption&gt;
&lt;p&gt;Can you see that weird unexpected variation? It is the weekend effect! (That&amp;rsquo;s why in the next dashboard version I will be highlighting the weekends). Source:&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Can you see how all figures suffer from a dramatic fall on 29th and 30th March just before peaking again the day after? This is indeed an unexpected behaviour and really difficult to explain. Unless we realise that those days were Saturday and Sunday, and due to the fact that there are fewer people working at hospitals, data is not taken as fast as usual and therefore, accumulates on Monday. This phenomenon has been called &amp;ldquo;the weekend effect&amp;rdquo; (Did I mention that you should not take data too seriously? &amp;#x1f609;)&lt;/p&gt;
&lt;p&gt;It is because of that that some others, such as
from Financial Times, prefer to use a rolling average of a fixed period (such as 3 or 5 days), which is a more stable figure, such as that shown in the figure below.&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/FT-rollingdata.png"
alt="Source:"&gt;&lt;figcaption&gt;
&lt;p&gt;Source:&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Other techniques used to fix those outliers and display tendencies are to use smooth line plots based on the actual data, like the following plot made by
.&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/numeroteca-smooth.jpeg"
alt="Smoothed lines based on actual data. Source:"&gt;&lt;figcaption&gt;
&lt;p&gt;Smoothed lines based on actual data. Source:&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="4-do-not-make-hasty-comparisons"&gt;4. Do not make hasty comparisons&lt;/h2&gt;
&lt;p&gt;Surely, the most frequent type of visual is that comparing how COVID-19 is affecting different regions, either within a country or comparing different countries (usually using China or Wuchan as a reference -after all, it is where it all started). While this kind of plots could provide answers to questions such as how a specific region is doing regarding another one (and therefore, replicating or avoiding their measures against COVID-19, for example), those comparisons are really problematic. For starters, the fact that population or size is largely different invalidates any comparison in absolute terms.&lt;/p&gt;
&lt;p&gt;But even when using relative values (eg: number of cases per inhabitant), there are other key factors that have a direct impact on the evolution of the disease and its effects are assumed to be the same, while in reality can differ in several orders of magnitude, such as demography&lt;sup id="fnref:8"&gt;&lt;a href="#fn:8" class="footnote-ref" role="doc-noteref"&gt;8&lt;/a&gt;&lt;/sup&gt;, geography, urban settlements or health systems (in terms of human and financial resources), just to name a few.&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/Diario_es_demografia_infectados2.png"
alt="Histogram of cases and mortality rates by age and country. Most cases in Italy were from 70-90, as opposed to 20-30 in South Korea, hence the enormous difference in the total of deaths. Source:"&gt;&lt;figcaption&gt;
&lt;p&gt;Histogram of cases and mortality rates by age and country. Most cases in Italy were from 70-90, as opposed to 20-30 in South Korea, hence the enormous difference in the total of deaths. Source:&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;strong&gt;Amongst all those differences, though, the most relevant well may be the number of tests done to detect COVID-19, a figure that is not just different but also usually unknown&lt;/strong&gt;. This is by no means trivial, as we have seen that the number of cases is defined using this parameter. Therefore, if a country performed a really low number of tests, it will also have an extremely low number of cases. Out of sight, out of mind.&lt;/p&gt;
&lt;p&gt;So, the biggest problem here is that without taking into consideration those factors, comparisons may render plenty of biased conclusions that have nothing to do with reality, such as come countries may be dodging COVID-19 or are kind of immune, or that COVID-19 only affects countries with a bad health system, unorganized governments, or undeveloped countries. As a result, &lt;strong&gt;there is a risk of developing a narrative of moral superiority&lt;sup id="fnref:9"&gt;&lt;a href="#fn:9" class="footnote-ref" role="doc-noteref"&gt;9&lt;/a&gt;&lt;/sup&gt; based on totally wrong foundations&lt;/strong&gt; like what some politicians have started to do in their own self-interest&lt;sup id="fnref:10"&gt;&lt;a href="#fn:10" class="footnote-ref" role="doc-noteref"&gt;10&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/repugnant-conclusions.jpg"
alt="Don&amp;rsquo;t be like Wopke Hoekstra doing hasty comparison, or you risk reaching repugnant conclusions like him. Photo:"&gt;&lt;figcaption&gt;
&lt;p&gt;Don&amp;rsquo;t be like Wopke Hoekstra doing hasty comparison, or you risk reaching repugnant conclusions like him. Photo:&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="5-take-predictions-even-less-seriously"&gt;5. Take predictions even less seriously&lt;/h2&gt;
&lt;p&gt;The last group of plots, and the most complex ones, are those that make predictions. While they are really appealing and, apparently, provide answers to one of our main concerns (&lt;em&gt;&amp;ldquo;When is this going to end?&amp;rdquo; / Will this last any longer?&lt;/em&gt;) in a very understandable way, they are really tricky. There are several ways to make predictions, such as using linear regression or models. While a model can be as easy&lt;sup id="fnref:11"&gt;&lt;a href="#fn:11" class="footnote-ref" role="doc-noteref"&gt;11&lt;/a&gt;&lt;/sup&gt; or as complex as we want it to be (and as a result, their accuracy will differ dramatically), they mostly rely on having a good set of historic data or knowing the logics of the phenomenon they want to describe. Unfortunately, since COVID-19 is a new phenomenon, we are lacking of both, and therefore, predictions at this stage are prone to errors. Some predictions are based on what has happened in other places where the outbreak started before, but we have seen how problematic this can be.&lt;/p&gt;
&lt;figure&gt;&lt;img src="https://carlos-hugoblox.netlify.app/en/blog/2020/04/considerations-on-the-importance-of-data-and-science-in-data-science/img/tarot.jpeg"
alt="If we can&amp;rsquo;t see (nor understand) the model behind a prediction, it can be as useful and reliable as that of a fortune teller."&gt;&lt;figcaption&gt;
&lt;p&gt;If we can&amp;rsquo;t see (nor understand) the model behind a prediction, it can be as useful and reliable as that of a fortune teller.&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Last, but not least, we should not forget that they require a large amount of knowledge of a particular field something like most of the people who are doing those nice visualizations (including myself) lack of. Therefore, since I have openly admitted that this is something beyond my knowledge, at this point I can only but recommend to be sceptical about any prediction that is not made by an authority on the field. And even in that case (provided I could understand it), I would recommend caution. Or even better: just rely on predictions if they were made by epidemiologists and you are one of them.&lt;/p&gt;
&lt;h2 id="wrapping-up"&gt;Wrapping up&lt;/h2&gt;
&lt;p&gt;As argued, if data visualization is never easy, it is even less so in the case of a novel phenomenon such as COVID-19. Therefore, when facing any type of visuals, we should proceed with caution. If you are doing (or planning to do) any type of visualization, ask yourself what question do you want to answer and which is the best way to do it, take into account the aforementioned considerations and make them evident to your readers&lt;sup id="fnref:12"&gt;&lt;a href="#fn:12" class="footnote-ref" role="doc-noteref"&gt;12&lt;/a&gt;&lt;/sup&gt;. Also, make your analysis reproducible, so anyone could tell you if you did something wrong or even fix it by themselves. If you are simply watching them, look for all those explanations, and if you can&amp;rsquo;t find them, ask for them, help the author or simply ignore it and look for an alternative. But in any case, you should always remember not to take data too seriously or too blindly. Data by itself is not what really matters, is what we do with it and how we do it in order to achieve knowledge what really matters. And here&amp;rsquo;s when science plays its role.&lt;/p&gt;
&lt;div class="footnotes" role="doc-endnotes"&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id="fn:1"&gt;
&lt;p&gt;There is a third scenario: to infer or calculate the data from other datasets.&amp;#160;&lt;a href="#fnref:1" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:2"&gt;
&lt;p&gt;For example,
, without a clean structure. Not only they use a proprietary format, but they mix data with metadata on the same file, even in the same sheet, instead of using a wide format where every column is a field and every row is an observation or a long format where keys and values are stored as in a dictionary. On the other hand, Spain decided to release the data in PDF. PDFs are great for visualization because they are an ISO format and can be opened with plenty of different softwares. Unfortunately, it is not great for consuming data, as it is not structured in any way.&amp;#160;&lt;a href="#fnref:2" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:3"&gt;
&lt;p&gt;At the time of writing this post, and according to
&amp;rsquo;s page, the
. The standard method of testing is real-time reverse transcription polymerase chain reaction (rRT-PCR), typically done on respiratory samples obtained by a nasopharyngeal swab and results are generally available within a few hours to two days.&amp;#160;&lt;a href="#fnref:3" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:4"&gt;
&lt;p&gt;This article from El Pais (in Spanish)
&amp;#160;&lt;a href="#fnref:4" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:5"&gt;
&lt;p&gt;
&amp;#160;&lt;a href="#fnref:5" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:6"&gt;
&lt;p&gt;This is the commonly phrased version of
principle.&amp;#160;&lt;a href="#fnref:6" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:7"&gt;
&lt;p&gt;In Spain
&amp;#160;&lt;a href="#fnref:7" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:8"&gt;
&lt;p&gt;Refer to this (
-in Spanish) which focuses on the fact that coronavirus mortality rate differs enormously according to the age.&amp;#160;&lt;a href="#fnref:8" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:9"&gt;
&lt;p&gt;Refer to
(Publico, in Spanish)&amp;#160;&lt;a href="#fnref:9" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:10"&gt;
&lt;p&gt;Take as an example the unfortunate words by Dutch finance minister Wopke Hoekstra, who suggested the EU &amp;ldquo;should investigate countries like Spain that say they have no budgetary margin to deal with the effects of the crisis provoked by the new coronavirus in spite of the fact that the eurozone has grown for seven consecutive years&amp;rdquo;, a statement that was later qualified as &amp;ldquo;repugnant&amp;rdquo; by Portugal&amp;rsquo;s Primer Minister, António Costa. (Source:
)&amp;#160;&lt;a href="#fnref:10" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:11"&gt;
&lt;p&gt;Just to give you an example of how easy can be to implement a prediction in R, refer to this article:
&amp;#160;&lt;a href="#fnref:11" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:12"&gt;
&lt;p&gt;Also, reading this post may be useful:
&amp;#160;&lt;a href="#fnref:12" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</description></item><item><title>Manipulating dataframes in R and Python</title><link>https://carlos-hugoblox.netlify.app/en/blog/2020/02/manipulating-dataframes-in-r-and-python/</link><pubDate>Sat, 08 Feb 2020 17:46:34 +0000</pubDate><guid>https://carlos-hugoblox.netlify.app/en/blog/2020/02/manipulating-dataframes-in-r-and-python/</guid><description>&lt;p&gt;While I have been using &lt;code&gt;R&lt;/code&gt; for many years now (mainly for data manipulation and visualization), and I am extremely happy with some of its features (like how easy is to deal with data or to create interactive reports that can be exported in plenty of different outputs, such as pdf, documents, slides, dashboards or blog posts like this one). However, I have always wanted to learn &lt;code&gt;python&lt;/code&gt;, mostly because it is a multi-purpose language that I would be able to use in other aspects of my everyday life such as web development, &lt;code&gt;QGIS&lt;/code&gt; or Academic research. It is for that reason that I have recently started to learn &lt;code&gt;python&lt;/code&gt;&amp;rsquo;s &lt;code&gt;pandas&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;In the following blog post I will be comparing how to perform the same tasks using &lt;code&gt;pandas&lt;/code&gt; and &lt;code&gt;tidyverse&lt;/code&gt;. This mainly serves two learning outcomes:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;To generate a cheat sheet that can work as a reminder (for myself): I know there are pages like
, but I learn by doing, so I needed to write the code myself)&lt;/li&gt;
&lt;li&gt;To use &lt;code&gt;reticulate&lt;/code&gt; package, which allows running both, &lt;code&gt;R&lt;/code&gt; and &lt;code&gt;python&lt;/code&gt; within the same document (a &lt;code&gt;Rmarkdown&lt;/code&gt; file to be more specific)&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="loading-environment"&gt;Loading environment&lt;/h2&gt;
&lt;p&gt;Since I want to use &lt;code&gt;Python&lt;/code&gt; and &lt;code&gt;R&lt;/code&gt; from a &lt;code&gt;.Rmarkdown&lt;/code&gt; file, I first need to load &lt;code&gt;reticulate&lt;/code&gt; for this, which is a &lt;code&gt;python&lt;/code&gt; interface for &lt;code&gt;R&lt;/code&gt;. Also, since &lt;code&gt;pandas&lt;/code&gt; is not a standard library module, I need to load a python environment&lt;sup id="fnref:1"&gt;&lt;a href="#fn:1" class="footnote-ref" role="doc-noteref"&gt;1&lt;/a&gt;&lt;/sup&gt; with the required packages. &lt;code&gt;reticulate&lt;/code&gt; makes it possible to load environments created with &lt;code&gt;anaconda&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;The last step is to load some data about COVID-19 provided by &lt;code&gt;coronavirus&lt;/code&gt; package, which I will be using in this blog post.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-r" data-lang="r"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nf"&gt;library&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;reticulate&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nf"&gt;library&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;DT&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nf"&gt;use_condaenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;&amp;#39;osm_imports_preparations&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;required&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;TRUE&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="loading-data"&gt;Loading data&lt;/h2&gt;
&lt;p&gt;First thing we are doing to do is to read a CSV file and turn it into a dataframe which we are going to manipulate in the next steps.&lt;/p&gt;
&lt;h3 id="r"&gt;R&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-r" data-lang="r"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# We will be loading a CSV file from RamiKrispin&amp;#39;s coronavirus&amp;#39; package.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;csv_url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s"&gt;&amp;#34;https://raw.githubusercontent.com/RamiKrispin/coronavirus/master/csv/coronavirus.csv&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Read the CSV, convert it into a dataframe and store it in a variable.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;r_df&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;-&lt;/span&gt; &lt;span class="nf"&gt;read.csv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;csv_url&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Explore the first on the dataframe.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nf"&gt;head&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r_df&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="m"&gt;12&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## date province country lat long type cases
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 1 2020-01-22 Afghanistan 33.93911 67.70995 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 2 2020-01-23 Afghanistan 33.93911 67.70995 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 3 2020-01-24 Afghanistan 33.93911 67.70995 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 4 2020-01-25 Afghanistan 33.93911 67.70995 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 5 2020-01-26 Afghanistan 33.93911 67.70995 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 6 2020-01-27 Afghanistan 33.93911 67.70995 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 7 2020-01-28 Afghanistan 33.93911 67.70995 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 8 2020-01-29 Afghanistan 33.93911 67.70995 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 9 2020-01-30 Afghanistan 33.93911 67.70995 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 10 2020-01-31 Afghanistan 33.93911 67.70995 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 11 2020-02-01 Afghanistan 33.93911 67.70995 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 12 2020-02-02 Afghanistan 33.93911 67.70995 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="python"&gt;Python&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;pandas&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nn"&gt;pd&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# We will be loading a CSV file from RamiKrispin&amp;#39;s coronavirus&amp;#39; package.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;csv_url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;https://raw.githubusercontent.com/RamiKrispin/coronavirus/master/csv/coronavirus.csv&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Read the CSV, convert it into a dataframe and store it in a variable.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;py_df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;read_csv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;csv_url&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Explore the first elements on the dataframe.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;py_df&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;head&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## date province country lat long type cases
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 0 2020-01-22 NaN Afghanistan 33.93911 67.709953 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 1 2020-01-23 NaN Afghanistan 33.93911 67.709953 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 2 2020-01-24 NaN Afghanistan 33.93911 67.709953 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 3 2020-01-25 NaN Afghanistan 33.93911 67.709953 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 4 2020-01-26 NaN Afghanistan 33.93911 67.709953 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 5 2020-01-27 NaN Afghanistan 33.93911 67.709953 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 6 2020-01-28 NaN Afghanistan 33.93911 67.709953 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 7 2020-01-29 NaN Afghanistan 33.93911 67.709953 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 8 2020-01-30 NaN Afghanistan 33.93911 67.709953 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 9 2020-01-31 NaN Afghanistan 33.93911 67.709953 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 10 2020-02-01 NaN Afghanistan 33.93911 67.709953 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 11 2020-02-02 NaN Afghanistan 33.93911 67.709953 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;So far, there are no significant differences between both, but note the following differences:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;code&gt;R&lt;/code&gt; works with regular &lt;strong&gt;functions&lt;/strong&gt;, whereas &lt;code&gt;python&lt;/code&gt; uses &lt;strong&gt;methods&lt;/strong&gt; instead.&lt;/li&gt;
&lt;li&gt;In order to work with dataframes in &lt;code&gt;python&lt;/code&gt;, &lt;code&gt;pandas&lt;/code&gt; module has to be imported beforehand, whereas it is a base feature from &lt;code&gt;R&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Both commands have read the same file but, apparently, they display different data. We will need to explore further the imported data to make sure that both are what we expected and, therefore, the same.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="exploring-a-dataframe"&gt;Exploring a dataframe&lt;/h2&gt;
&lt;p&gt;In this step we are going to evaluate what kind of object have we created, as well as a very basic data exploration.&lt;/p&gt;
&lt;h3 id="r-1"&gt;R&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-r" data-lang="r"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Explore what kind of object r_df is.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r_df&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-gdscript3" data-lang="gdscript3"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;## &amp;#39;data.frame&amp;#39;: 218276 obs. of 7 variables:&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;## $ date : chr &amp;#34;2020-01-22&amp;#34; &amp;#34;2020-01-23&amp;#34; &amp;#34;2020-01-24&amp;#34; &amp;#34;2020-01-25&amp;#34; ...&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;## $ province: chr &amp;#34;&amp;#34; &amp;#34;&amp;#34; &amp;#34;&amp;#34; &amp;#34;&amp;#34; ...&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;## $ country : chr &amp;#34;Afghanistan&amp;#34; &amp;#34;Afghanistan&amp;#34; &amp;#34;Afghanistan&amp;#34; &amp;#34;Afghanistan&amp;#34; ...&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;## $ lat : num 33.9 33.9 33.9 33.9 33.9 ...&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;## $ long : num 67.7 67.7 67.7 67.7 67.7 ...&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;## $ type : chr &amp;#34;confirmed&amp;#34; &amp;#34;confirmed&amp;#34; &amp;#34;confirmed&amp;#34; &amp;#34;confirmed&amp;#34; ...&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;## $ cases : int 0 0 0 0 0 0 0 0 0 0 ...&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-r" data-lang="r"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Basic statistics&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nf"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r_df&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## date province country lat
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Length:218276 Length:218276 Length:218276 Min. :-51.796
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Class :character Class :character Class :character 1st Qu.: 6.428
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Mode :character Mode :character Mode :character Median : 22.041
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Mean : 20.561
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 3rd Qu.: 40.182
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Max. : 71.707
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## long type cases
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Min. :-135.00 Length:218276 Min. :-16298.0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 1st Qu.: -12.89 Class :character 1st Qu.: 0.0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Median : 21.75 Mode :character Median : 0.0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Mean : 25.01 Mean : 331.8
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 3rd Qu.: 84.25 3rd Qu.: 12.0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Max. : 178.06 Max. :140050.0
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-r" data-lang="r"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Explore unique values within a variable.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nf"&gt;levels&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;as.factor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r_df&lt;/span&gt;&lt;span class="o"&gt;$&lt;/span&gt;&lt;span class="n"&gt;date&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [1] &amp;#34;2020-01-22&amp;#34; &amp;#34;2020-01-23&amp;#34; &amp;#34;2020-01-24&amp;#34; &amp;#34;2020-01-25&amp;#34; &amp;#34;2020-01-26&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [6] &amp;#34;2020-01-27&amp;#34; &amp;#34;2020-01-28&amp;#34; &amp;#34;2020-01-29&amp;#34; &amp;#34;2020-01-30&amp;#34; &amp;#34;2020-01-31&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [11] &amp;#34;2020-02-01&amp;#34; &amp;#34;2020-02-02&amp;#34; &amp;#34;2020-02-03&amp;#34; &amp;#34;2020-02-04&amp;#34; &amp;#34;2020-02-05&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [16] &amp;#34;2020-02-06&amp;#34; &amp;#34;2020-02-07&amp;#34; &amp;#34;2020-02-08&amp;#34; &amp;#34;2020-02-09&amp;#34; &amp;#34;2020-02-10&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [21] &amp;#34;2020-02-11&amp;#34; &amp;#34;2020-02-12&amp;#34; &amp;#34;2020-02-13&amp;#34; &amp;#34;2020-02-14&amp;#34; &amp;#34;2020-02-15&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [26] &amp;#34;2020-02-16&amp;#34; &amp;#34;2020-02-17&amp;#34; &amp;#34;2020-02-18&amp;#34; &amp;#34;2020-02-19&amp;#34; &amp;#34;2020-02-20&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [31] &amp;#34;2020-02-21&amp;#34; &amp;#34;2020-02-22&amp;#34; &amp;#34;2020-02-23&amp;#34; &amp;#34;2020-02-24&amp;#34; &amp;#34;2020-02-25&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [36] &amp;#34;2020-02-26&amp;#34; &amp;#34;2020-02-27&amp;#34; &amp;#34;2020-02-28&amp;#34; &amp;#34;2020-02-29&amp;#34; &amp;#34;2020-03-01&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [41] &amp;#34;2020-03-02&amp;#34; &amp;#34;2020-03-03&amp;#34; &amp;#34;2020-03-04&amp;#34; &amp;#34;2020-03-05&amp;#34; &amp;#34;2020-03-06&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [46] &amp;#34;2020-03-07&amp;#34; &amp;#34;2020-03-08&amp;#34; &amp;#34;2020-03-09&amp;#34; &amp;#34;2020-03-10&amp;#34; &amp;#34;2020-03-11&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [51] &amp;#34;2020-03-12&amp;#34; &amp;#34;2020-03-13&amp;#34; &amp;#34;2020-03-14&amp;#34; &amp;#34;2020-03-15&amp;#34; &amp;#34;2020-03-16&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [56] &amp;#34;2020-03-17&amp;#34; &amp;#34;2020-03-18&amp;#34; &amp;#34;2020-03-19&amp;#34; &amp;#34;2020-03-20&amp;#34; &amp;#34;2020-03-21&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [61] &amp;#34;2020-03-22&amp;#34; &amp;#34;2020-03-23&amp;#34; &amp;#34;2020-03-24&amp;#34; &amp;#34;2020-03-25&amp;#34; &amp;#34;2020-03-26&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [66] &amp;#34;2020-03-27&amp;#34; &amp;#34;2020-03-28&amp;#34; &amp;#34;2020-03-29&amp;#34; &amp;#34;2020-03-30&amp;#34; &amp;#34;2020-03-31&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [71] &amp;#34;2020-04-01&amp;#34; &amp;#34;2020-04-02&amp;#34; &amp;#34;2020-04-03&amp;#34; &amp;#34;2020-04-04&amp;#34; &amp;#34;2020-04-05&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [76] &amp;#34;2020-04-06&amp;#34; &amp;#34;2020-04-07&amp;#34; &amp;#34;2020-04-08&amp;#34; &amp;#34;2020-04-09&amp;#34; &amp;#34;2020-04-10&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [81] &amp;#34;2020-04-11&amp;#34; &amp;#34;2020-04-12&amp;#34; &amp;#34;2020-04-13&amp;#34; &amp;#34;2020-04-14&amp;#34; &amp;#34;2020-04-15&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [86] &amp;#34;2020-04-16&amp;#34; &amp;#34;2020-04-17&amp;#34; &amp;#34;2020-04-18&amp;#34; &amp;#34;2020-04-19&amp;#34; &amp;#34;2020-04-20&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [91] &amp;#34;2020-04-21&amp;#34; &amp;#34;2020-04-22&amp;#34; &amp;#34;2020-04-23&amp;#34; &amp;#34;2020-04-24&amp;#34; &amp;#34;2020-04-25&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [96] &amp;#34;2020-04-26&amp;#34; &amp;#34;2020-04-27&amp;#34; &amp;#34;2020-04-28&amp;#34; &amp;#34;2020-04-29&amp;#34; &amp;#34;2020-04-30&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [101] &amp;#34;2020-05-01&amp;#34; &amp;#34;2020-05-02&amp;#34; &amp;#34;2020-05-03&amp;#34; &amp;#34;2020-05-04&amp;#34; &amp;#34;2020-05-05&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [106] &amp;#34;2020-05-06&amp;#34; &amp;#34;2020-05-07&amp;#34; &amp;#34;2020-05-08&amp;#34; &amp;#34;2020-05-09&amp;#34; &amp;#34;2020-05-10&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [111] &amp;#34;2020-05-11&amp;#34; &amp;#34;2020-05-12&amp;#34; &amp;#34;2020-05-13&amp;#34; &amp;#34;2020-05-14&amp;#34; &amp;#34;2020-05-15&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [116] &amp;#34;2020-05-16&amp;#34; &amp;#34;2020-05-17&amp;#34; &amp;#34;2020-05-18&amp;#34; &amp;#34;2020-05-19&amp;#34; &amp;#34;2020-05-20&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [121] &amp;#34;2020-05-21&amp;#34; &amp;#34;2020-05-22&amp;#34; &amp;#34;2020-05-23&amp;#34; &amp;#34;2020-05-24&amp;#34; &amp;#34;2020-05-25&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [126] &amp;#34;2020-05-26&amp;#34; &amp;#34;2020-05-27&amp;#34; &amp;#34;2020-05-28&amp;#34; &amp;#34;2020-05-29&amp;#34; &amp;#34;2020-05-30&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [131] &amp;#34;2020-05-31&amp;#34; &amp;#34;2020-06-01&amp;#34; &amp;#34;2020-06-02&amp;#34; &amp;#34;2020-06-03&amp;#34; &amp;#34;2020-06-04&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [136] &amp;#34;2020-06-05&amp;#34; &amp;#34;2020-06-06&amp;#34; &amp;#34;2020-06-07&amp;#34; &amp;#34;2020-06-08&amp;#34; &amp;#34;2020-06-09&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [141] &amp;#34;2020-06-10&amp;#34; &amp;#34;2020-06-11&amp;#34; &amp;#34;2020-06-12&amp;#34; &amp;#34;2020-06-13&amp;#34; &amp;#34;2020-06-14&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [146] &amp;#34;2020-06-15&amp;#34; &amp;#34;2020-06-16&amp;#34; &amp;#34;2020-06-17&amp;#34; &amp;#34;2020-06-18&amp;#34; &amp;#34;2020-06-19&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [151] &amp;#34;2020-06-20&amp;#34; &amp;#34;2020-06-21&amp;#34; &amp;#34;2020-06-22&amp;#34; &amp;#34;2020-06-23&amp;#34; &amp;#34;2020-06-24&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [156] &amp;#34;2020-06-25&amp;#34; &amp;#34;2020-06-26&amp;#34; &amp;#34;2020-06-27&amp;#34; &amp;#34;2020-06-28&amp;#34; &amp;#34;2020-06-29&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [161] &amp;#34;2020-06-30&amp;#34; &amp;#34;2020-07-01&amp;#34; &amp;#34;2020-07-02&amp;#34; &amp;#34;2020-07-03&amp;#34; &amp;#34;2020-07-04&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [166] &amp;#34;2020-07-05&amp;#34; &amp;#34;2020-07-06&amp;#34; &amp;#34;2020-07-07&amp;#34; &amp;#34;2020-07-08&amp;#34; &amp;#34;2020-07-09&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [171] &amp;#34;2020-07-10&amp;#34; &amp;#34;2020-07-11&amp;#34; &amp;#34;2020-07-12&amp;#34; &amp;#34;2020-07-13&amp;#34; &amp;#34;2020-07-14&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [176] &amp;#34;2020-07-15&amp;#34; &amp;#34;2020-07-16&amp;#34; &amp;#34;2020-07-17&amp;#34; &amp;#34;2020-07-18&amp;#34; &amp;#34;2020-07-19&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [181] &amp;#34;2020-07-20&amp;#34; &amp;#34;2020-07-21&amp;#34; &amp;#34;2020-07-22&amp;#34; &amp;#34;2020-07-23&amp;#34; &amp;#34;2020-07-24&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [186] &amp;#34;2020-07-25&amp;#34; &amp;#34;2020-07-26&amp;#34; &amp;#34;2020-07-27&amp;#34; &amp;#34;2020-07-28&amp;#34; &amp;#34;2020-07-29&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [191] &amp;#34;2020-07-30&amp;#34; &amp;#34;2020-07-31&amp;#34; &amp;#34;2020-08-01&amp;#34; &amp;#34;2020-08-02&amp;#34; &amp;#34;2020-08-03&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [196] &amp;#34;2020-08-04&amp;#34; &amp;#34;2020-08-05&amp;#34; &amp;#34;2020-08-06&amp;#34; &amp;#34;2020-08-07&amp;#34; &amp;#34;2020-08-08&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [201] &amp;#34;2020-08-09&amp;#34; &amp;#34;2020-08-10&amp;#34; &amp;#34;2020-08-11&amp;#34; &amp;#34;2020-08-12&amp;#34; &amp;#34;2020-08-13&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [206] &amp;#34;2020-08-14&amp;#34; &amp;#34;2020-08-15&amp;#34; &amp;#34;2020-08-16&amp;#34; &amp;#34;2020-08-17&amp;#34; &amp;#34;2020-08-18&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [211] &amp;#34;2020-08-19&amp;#34; &amp;#34;2020-08-20&amp;#34; &amp;#34;2020-08-21&amp;#34; &amp;#34;2020-08-22&amp;#34; &amp;#34;2020-08-23&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [216] &amp;#34;2020-08-24&amp;#34; &amp;#34;2020-08-25&amp;#34; &amp;#34;2020-08-26&amp;#34; &amp;#34;2020-08-27&amp;#34; &amp;#34;2020-08-28&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [221] &amp;#34;2020-08-29&amp;#34; &amp;#34;2020-08-30&amp;#34; &amp;#34;2020-08-31&amp;#34; &amp;#34;2020-09-01&amp;#34; &amp;#34;2020-09-02&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [226] &amp;#34;2020-09-03&amp;#34; &amp;#34;2020-09-04&amp;#34; &amp;#34;2020-09-05&amp;#34; &amp;#34;2020-09-06&amp;#34; &amp;#34;2020-09-07&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [231] &amp;#34;2020-09-08&amp;#34; &amp;#34;2020-09-09&amp;#34; &amp;#34;2020-09-10&amp;#34; &amp;#34;2020-09-11&amp;#34; &amp;#34;2020-09-12&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [236] &amp;#34;2020-09-13&amp;#34; &amp;#34;2020-09-14&amp;#34; &amp;#34;2020-09-15&amp;#34; &amp;#34;2020-09-16&amp;#34; &amp;#34;2020-09-17&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [241] &amp;#34;2020-09-18&amp;#34; &amp;#34;2020-09-19&amp;#34; &amp;#34;2020-09-20&amp;#34; &amp;#34;2020-09-21&amp;#34; &amp;#34;2020-09-22&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [246] &amp;#34;2020-09-23&amp;#34; &amp;#34;2020-09-24&amp;#34; &amp;#34;2020-09-25&amp;#34; &amp;#34;2020-09-26&amp;#34; &amp;#34;2020-09-27&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [251] &amp;#34;2020-09-28&amp;#34; &amp;#34;2020-09-29&amp;#34; &amp;#34;2020-09-30&amp;#34; &amp;#34;2020-10-01&amp;#34; &amp;#34;2020-10-02&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [256] &amp;#34;2020-10-03&amp;#34; &amp;#34;2020-10-04&amp;#34; &amp;#34;2020-10-05&amp;#34; &amp;#34;2020-10-06&amp;#34; &amp;#34;2020-10-07&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [261] &amp;#34;2020-10-08&amp;#34; &amp;#34;2020-10-09&amp;#34; &amp;#34;2020-10-10&amp;#34; &amp;#34;2020-10-11&amp;#34; &amp;#34;2020-10-12&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [266] &amp;#34;2020-10-13&amp;#34; &amp;#34;2020-10-14&amp;#34; &amp;#34;2020-10-15&amp;#34; &amp;#34;2020-10-16&amp;#34; &amp;#34;2020-10-17&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [271] &amp;#34;2020-10-18&amp;#34; &amp;#34;2020-10-19&amp;#34; &amp;#34;2020-10-20&amp;#34; &amp;#34;2020-10-21&amp;#34; &amp;#34;2020-10-22&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## [276] &amp;#34;2020-10-23&amp;#34; &amp;#34;2020-10-24&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="python-1"&gt;Python&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Explore what kind of entity py_df is.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;py_df&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;info&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Now, calculate basic statistics for the numeric columns in the DataFrame.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;lt;class &amp;#39;pandas.core.frame.DataFrame&amp;#39;&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## RangeIndex: 218276 entries, 0 to 218275
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Data columns (total 7 columns):
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## # Column Non-Null Count Dtype
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## --- ------ -------------- -----
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 0 date 218276 non-null object
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 1 province 63433 non-null object
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 2 country 218276 non-null object
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 3 lat 218276 non-null float64
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 4 long 218276 non-null float64
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 5 type 218276 non-null object
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 6 cases 218276 non-null int64
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## dtypes: float64(2), int64(1), object(4)
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## memory usage: 11.7+ MB
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;py_df&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# List unique values in the df[&amp;#39;date&amp;#39;] column&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## lat long cases
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## count 218276.000000 218276.000000 218276.000000
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## mean 20.561062 25.011408 331.749812
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## std 24.759252 69.572388 3045.547043
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## min -51.796300 -135.000000 -16298.000000
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 25% 6.428055 -12.885800 0.000000
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 50% 22.041450 21.745300 0.000000
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 75% 40.182400 84.250000 12.000000
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## max 71.706900 178.065000 140050.000000
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;py_df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;date&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;unique&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="c1"&gt;# I prefer using this notation to prevent problems with columns with a dot inside.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## array([&amp;#39;2020-01-22&amp;#39;, &amp;#39;2020-01-23&amp;#39;, &amp;#39;2020-01-24&amp;#39;, &amp;#39;2020-01-25&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-01-26&amp;#39;, &amp;#39;2020-01-27&amp;#39;, &amp;#39;2020-01-28&amp;#39;, &amp;#39;2020-01-29&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-01-30&amp;#39;, &amp;#39;2020-01-31&amp;#39;, &amp;#39;2020-02-01&amp;#39;, &amp;#39;2020-02-02&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-02-03&amp;#39;, &amp;#39;2020-02-04&amp;#39;, &amp;#39;2020-02-05&amp;#39;, &amp;#39;2020-02-06&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-02-07&amp;#39;, &amp;#39;2020-02-08&amp;#39;, &amp;#39;2020-02-09&amp;#39;, &amp;#39;2020-02-10&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-02-11&amp;#39;, &amp;#39;2020-02-12&amp;#39;, &amp;#39;2020-02-13&amp;#39;, &amp;#39;2020-02-14&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-02-15&amp;#39;, &amp;#39;2020-02-16&amp;#39;, &amp;#39;2020-02-17&amp;#39;, &amp;#39;2020-02-18&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-02-19&amp;#39;, &amp;#39;2020-02-20&amp;#39;, &amp;#39;2020-02-21&amp;#39;, &amp;#39;2020-02-22&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-02-23&amp;#39;, &amp;#39;2020-02-24&amp;#39;, &amp;#39;2020-02-25&amp;#39;, &amp;#39;2020-02-26&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-02-27&amp;#39;, &amp;#39;2020-02-28&amp;#39;, &amp;#39;2020-02-29&amp;#39;, &amp;#39;2020-03-01&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-03-02&amp;#39;, &amp;#39;2020-03-03&amp;#39;, &amp;#39;2020-03-04&amp;#39;, &amp;#39;2020-03-05&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-03-06&amp;#39;, &amp;#39;2020-03-07&amp;#39;, &amp;#39;2020-03-08&amp;#39;, &amp;#39;2020-03-09&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-03-10&amp;#39;, &amp;#39;2020-03-11&amp;#39;, &amp;#39;2020-03-12&amp;#39;, &amp;#39;2020-03-13&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-03-14&amp;#39;, &amp;#39;2020-03-15&amp;#39;, &amp;#39;2020-03-16&amp;#39;, &amp;#39;2020-03-17&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-03-18&amp;#39;, &amp;#39;2020-03-19&amp;#39;, &amp;#39;2020-03-20&amp;#39;, &amp;#39;2020-03-21&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-03-22&amp;#39;, &amp;#39;2020-03-23&amp;#39;, &amp;#39;2020-03-24&amp;#39;, &amp;#39;2020-03-25&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-03-26&amp;#39;, &amp;#39;2020-03-27&amp;#39;, &amp;#39;2020-03-28&amp;#39;, &amp;#39;2020-03-29&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-03-30&amp;#39;, &amp;#39;2020-03-31&amp;#39;, &amp;#39;2020-04-01&amp;#39;, &amp;#39;2020-04-02&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-04-03&amp;#39;, &amp;#39;2020-04-04&amp;#39;, &amp;#39;2020-04-05&amp;#39;, &amp;#39;2020-04-06&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-04-07&amp;#39;, &amp;#39;2020-04-08&amp;#39;, &amp;#39;2020-04-09&amp;#39;, &amp;#39;2020-04-10&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-04-11&amp;#39;, &amp;#39;2020-04-12&amp;#39;, &amp;#39;2020-04-13&amp;#39;, &amp;#39;2020-04-14&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-04-15&amp;#39;, &amp;#39;2020-04-16&amp;#39;, &amp;#39;2020-04-17&amp;#39;, &amp;#39;2020-04-18&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-04-19&amp;#39;, &amp;#39;2020-04-20&amp;#39;, &amp;#39;2020-04-21&amp;#39;, &amp;#39;2020-04-22&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-04-23&amp;#39;, &amp;#39;2020-04-24&amp;#39;, &amp;#39;2020-04-25&amp;#39;, &amp;#39;2020-04-26&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-04-27&amp;#39;, &amp;#39;2020-04-28&amp;#39;, &amp;#39;2020-04-29&amp;#39;, &amp;#39;2020-04-30&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-05-01&amp;#39;, &amp;#39;2020-05-02&amp;#39;, &amp;#39;2020-05-03&amp;#39;, &amp;#39;2020-05-04&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-05-05&amp;#39;, &amp;#39;2020-05-06&amp;#39;, &amp;#39;2020-05-07&amp;#39;, &amp;#39;2020-05-08&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-05-09&amp;#39;, &amp;#39;2020-05-10&amp;#39;, &amp;#39;2020-05-11&amp;#39;, &amp;#39;2020-05-12&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-05-13&amp;#39;, &amp;#39;2020-05-14&amp;#39;, &amp;#39;2020-05-15&amp;#39;, &amp;#39;2020-05-16&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-05-17&amp;#39;, &amp;#39;2020-05-18&amp;#39;, &amp;#39;2020-05-19&amp;#39;, &amp;#39;2020-05-20&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-05-21&amp;#39;, &amp;#39;2020-05-22&amp;#39;, &amp;#39;2020-05-23&amp;#39;, &amp;#39;2020-05-24&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-05-25&amp;#39;, &amp;#39;2020-05-26&amp;#39;, &amp;#39;2020-05-27&amp;#39;, &amp;#39;2020-05-28&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-05-29&amp;#39;, &amp;#39;2020-05-30&amp;#39;, &amp;#39;2020-05-31&amp;#39;, &amp;#39;2020-06-01&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-06-02&amp;#39;, &amp;#39;2020-06-03&amp;#39;, &amp;#39;2020-06-04&amp;#39;, &amp;#39;2020-06-05&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-06-06&amp;#39;, &amp;#39;2020-06-07&amp;#39;, &amp;#39;2020-06-08&amp;#39;, &amp;#39;2020-06-09&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-06-10&amp;#39;, &amp;#39;2020-06-11&amp;#39;, &amp;#39;2020-06-12&amp;#39;, &amp;#39;2020-06-13&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-06-14&amp;#39;, &amp;#39;2020-06-15&amp;#39;, &amp;#39;2020-06-16&amp;#39;, &amp;#39;2020-06-17&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-06-18&amp;#39;, &amp;#39;2020-06-19&amp;#39;, &amp;#39;2020-06-20&amp;#39;, &amp;#39;2020-06-21&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-06-22&amp;#39;, &amp;#39;2020-06-23&amp;#39;, &amp;#39;2020-06-24&amp;#39;, &amp;#39;2020-06-25&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-06-26&amp;#39;, &amp;#39;2020-06-27&amp;#39;, &amp;#39;2020-06-28&amp;#39;, &amp;#39;2020-06-29&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-06-30&amp;#39;, &amp;#39;2020-07-01&amp;#39;, &amp;#39;2020-07-02&amp;#39;, &amp;#39;2020-07-03&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-07-04&amp;#39;, &amp;#39;2020-07-05&amp;#39;, &amp;#39;2020-07-06&amp;#39;, &amp;#39;2020-07-07&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-07-08&amp;#39;, &amp;#39;2020-07-09&amp;#39;, &amp;#39;2020-07-10&amp;#39;, &amp;#39;2020-07-11&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-07-12&amp;#39;, &amp;#39;2020-07-13&amp;#39;, &amp;#39;2020-07-14&amp;#39;, &amp;#39;2020-07-15&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-07-16&amp;#39;, &amp;#39;2020-07-17&amp;#39;, &amp;#39;2020-07-18&amp;#39;, &amp;#39;2020-07-19&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-07-20&amp;#39;, &amp;#39;2020-07-21&amp;#39;, &amp;#39;2020-07-22&amp;#39;, &amp;#39;2020-07-23&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-07-24&amp;#39;, &amp;#39;2020-07-25&amp;#39;, &amp;#39;2020-07-26&amp;#39;, &amp;#39;2020-07-27&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-07-28&amp;#39;, &amp;#39;2020-07-29&amp;#39;, &amp;#39;2020-07-30&amp;#39;, &amp;#39;2020-07-31&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-08-01&amp;#39;, &amp;#39;2020-08-02&amp;#39;, &amp;#39;2020-08-03&amp;#39;, &amp;#39;2020-08-04&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-08-05&amp;#39;, &amp;#39;2020-08-06&amp;#39;, &amp;#39;2020-08-07&amp;#39;, &amp;#39;2020-08-08&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-08-09&amp;#39;, &amp;#39;2020-08-10&amp;#39;, &amp;#39;2020-08-11&amp;#39;, &amp;#39;2020-08-12&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-08-13&amp;#39;, &amp;#39;2020-08-14&amp;#39;, &amp;#39;2020-08-15&amp;#39;, &amp;#39;2020-08-16&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-08-17&amp;#39;, &amp;#39;2020-08-18&amp;#39;, &amp;#39;2020-08-19&amp;#39;, &amp;#39;2020-08-20&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-08-21&amp;#39;, &amp;#39;2020-08-22&amp;#39;, &amp;#39;2020-08-23&amp;#39;, &amp;#39;2020-08-24&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-08-25&amp;#39;, &amp;#39;2020-08-26&amp;#39;, &amp;#39;2020-08-27&amp;#39;, &amp;#39;2020-08-28&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-08-29&amp;#39;, &amp;#39;2020-08-30&amp;#39;, &amp;#39;2020-08-31&amp;#39;, &amp;#39;2020-09-01&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-09-02&amp;#39;, &amp;#39;2020-09-03&amp;#39;, &amp;#39;2020-09-04&amp;#39;, &amp;#39;2020-09-05&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-09-06&amp;#39;, &amp;#39;2020-09-07&amp;#39;, &amp;#39;2020-09-08&amp;#39;, &amp;#39;2020-09-09&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-09-10&amp;#39;, &amp;#39;2020-09-11&amp;#39;, &amp;#39;2020-09-12&amp;#39;, &amp;#39;2020-09-13&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-09-14&amp;#39;, &amp;#39;2020-09-15&amp;#39;, &amp;#39;2020-09-16&amp;#39;, &amp;#39;2020-09-17&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-09-18&amp;#39;, &amp;#39;2020-09-19&amp;#39;, &amp;#39;2020-09-20&amp;#39;, &amp;#39;2020-09-21&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-09-22&amp;#39;, &amp;#39;2020-09-23&amp;#39;, &amp;#39;2020-09-24&amp;#39;, &amp;#39;2020-09-25&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-09-26&amp;#39;, &amp;#39;2020-09-27&amp;#39;, &amp;#39;2020-09-28&amp;#39;, &amp;#39;2020-09-29&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-09-30&amp;#39;, &amp;#39;2020-10-01&amp;#39;, &amp;#39;2020-10-02&amp;#39;, &amp;#39;2020-10-03&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-10-04&amp;#39;, &amp;#39;2020-10-05&amp;#39;, &amp;#39;2020-10-06&amp;#39;, &amp;#39;2020-10-07&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-10-08&amp;#39;, &amp;#39;2020-10-09&amp;#39;, &amp;#39;2020-10-10&amp;#39;, &amp;#39;2020-10-11&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-10-12&amp;#39;, &amp;#39;2020-10-13&amp;#39;, &amp;#39;2020-10-14&amp;#39;, &amp;#39;2020-10-15&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-10-16&amp;#39;, &amp;#39;2020-10-17&amp;#39;, &amp;#39;2020-10-18&amp;#39;, &amp;#39;2020-10-19&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-10-20&amp;#39;, &amp;#39;2020-10-21&amp;#39;, &amp;#39;2020-10-22&amp;#39;, &amp;#39;2020-10-23&amp;#39;,
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;#39;2020-10-24&amp;#39;], dtype=object)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;There are no significant differences, nor in the syntax nor in the ouput.&lt;/p&gt;
&lt;p&gt;Note:&lt;/p&gt;
&lt;p&gt;We prefer to use this notation in order to prevent problems with columns with a dot inside.&lt;/p&gt;
&lt;h2 id="sorting-dataframe"&gt;Sorting dataframe&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-r" data-lang="r"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nf"&gt;library&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tidyverse&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## ── Attaching packages ─────────────────────────────────────── tidyverse 1.3.0 ──
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## ✓ ggplot2 3.3.2 ✓ purrr 0.3.4
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## ✓ tibble 3.0.4 ✓ dplyr 1.0.2
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## ✓ tidyr 1.1.2 ✓ stringr 1.4.0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## ✓ readr 1.3.1 ✓ forcats 0.5.0
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## x dplyr::filter() masks stats::filter()
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## x dplyr::lag() masks stats::lag()
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-r" data-lang="r"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nf"&gt;head&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;arrange&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r_df&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;country&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## date province country lat long type cases
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 1 2020-01-22 Afghanistan 33.93911 67.70995 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 2 2020-01-23 Afghanistan 33.93911 67.70995 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 3 2020-01-24 Afghanistan 33.93911 67.70995 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 4 2020-01-25 Afghanistan 33.93911 67.70995 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 5 2020-01-26 Afghanistan 33.93911 67.70995 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 6 2020-01-27 Afghanistan 33.93911 67.70995 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;py_df&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sort_values&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;country&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;head&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## date province country lat long type cases
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 0 2020-01-22 NaN Afghanistan 33.93911 67.709953 confirmed 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 147922 2020-01-26 NaN Afghanistan 33.93911 67.709953 recovered 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 147921 2020-01-25 NaN Afghanistan 33.93911 67.709953 recovered 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 147920 2020-01-24 NaN Afghanistan 33.93911 67.709953 recovered 0
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 147919 2020-01-23 NaN Afghanistan 33.93911 67.709953 recovered 0
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="select-and-summarise"&gt;Select and summarise&lt;/h2&gt;
&lt;p&gt;Let&amp;rsquo;s pretend that we want to have a table displaying the top 10 countries with the most number of confirmed cases until today (2020-10-25).&lt;/p&gt;
&lt;h3 id="r-2"&gt;R&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-r" data-lang="r"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;r_df&lt;/span&gt; &lt;span class="o"&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;# select(Country.Region, cases, type) %&amp;gt;% &lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;type&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="s"&gt;&amp;#34;confirmed&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nf"&gt;group_by&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;country&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nf"&gt;summarise&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cases&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="o"&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nf"&gt;arrange&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;desc&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;total&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="o"&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nf"&gt;head&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="m"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## # A tibble: 5 x 2
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## country total
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## &amp;lt;chr&amp;gt; &amp;lt;int&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 1 US 8575177
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 2 India 7814682
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 3 Brazil 5380635
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 4 Russia 1487260
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 5 France 1084659
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-r" data-lang="r"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Or, even more succintly:&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;r_df&lt;/span&gt; &lt;span class="o"&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;type&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="s"&gt;&amp;#34;confirmed&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nf"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;country&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;wt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cases&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sort&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;TRUE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s"&gt;&amp;#34;total&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nf"&gt;head&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="m"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## country total
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 1 US 8575177
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 2 India 7814682
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 3 Brazil 5380635
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 4 Russia 1487260
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 5 France 1084659
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="python-2"&gt;Python&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;py_df&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;loc&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;py_df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;type&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;confirmed&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;groupby&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;country&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;agg&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;cases&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;sum&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;rename&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;columns&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;cases&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;total&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sort_values&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;total&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;ascending&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;head&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## total
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## country
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## US 8575177
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## India 7814682
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Brazil 5380635
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Russia 1487260
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## France 1084659
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This is something funny. &lt;code&gt;python&lt;/code&gt; takes pride in stating that is an elegant and easy to read syntax due to its strict indentation syntax. On the other hand, I have met several &lt;code&gt;python&lt;/code&gt; proponents mocking about &lt;code&gt;R&lt;/code&gt;&amp;rsquo;s code to be quite obscure and difficult to memorise. While I often share the same views (especially when dealing with base &lt;code&gt;R&lt;/code&gt;&amp;rsquo;s syntax), I particularly find &lt;code&gt;tidypverse&lt;/code&gt;&amp;rsquo;s syntax far easier to read and memorise than &lt;code&gt;pandas&lt;/code&gt;&amp;rsquo; . While the first makes use of the pipe operator ( &lt;code&gt;%&amp;gt;%&lt;/code&gt;) to chain commands while preventing typing unnecessary data, the latter requires to concatenate up to six different methods in a single line, which becomes too long to read (and thus, not liked very much by
)&lt;/p&gt;
&lt;h2 id="joins-and-calculations"&gt;Joins and calculations&lt;/h2&gt;
&lt;p&gt;But that&amp;rsquo;s not fair, we are comparing countries with very different number of population! If we are to compare them, we need to use relative values. For example, we would need to create a ranking based on the total number of cases per 1000 habitants.&lt;/p&gt;
&lt;p&gt;In order to do so, we will need to do the following steps:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Load a dataframe with population data per each country&lt;/li&gt;
&lt;li&gt;Add the population data by join our existing dataframes with the newly created one in the previous step (left join)&lt;/li&gt;
&lt;li&gt;Calculate relative number of confirmed cases like this: &lt;code&gt;\(confirmed~rel = \frac{confirmed}{population}\)&lt;/code&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="r-3"&gt;R&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-r" data-lang="r"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Load population data.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;countries19&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;-&lt;/span&gt; &lt;span class="nf"&gt;read.csv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;data/countries_pop19.csv&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;r_df&lt;/span&gt; &lt;span class="o"&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;type&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="s"&gt;&amp;#34;confirmed&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nf"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;country&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;wt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cases&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sort&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;TRUE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s"&gt;&amp;#34;confirmed&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;# Add population column.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nf"&gt;left_join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;countries19&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;by&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;c&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;country&amp;#34;&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s"&gt;&amp;#34;Location&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="o"&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;# Calculate relative cases.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nf"&gt;mutate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;confirmed_rel&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;confirmed&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;PopTotal&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;# Select certain columns only.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nf"&gt;select&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;country&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;confirmed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;confirmed_rel&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;# Sort by confirmed_rel on descending order.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nf"&gt;arrange&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;desc&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;confirmed_rel&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="o"&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;# Display everything on a nice datatable.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nf"&gt;head&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="m"&gt;10&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## country confirmed confirmed_rel
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 1 Andorra 4038 52.34231
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 2 Bahrain 79975 48.73066
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 3 Qatar 130965 46.24354
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 4 Israel 309413 36.31875
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 5 Holy See 27 33.12883
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 6 Panama 128515 30.26417
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 7 Kuwait 120927 28.74371
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 8 Peru 883116 27.16406
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 9 Montenegro 16629 26.47981
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## 10 Belgium 305409 26.46680
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="python-3"&gt;Python&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Load population data.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;countries19&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;read_csv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;data/countries_pop19.csv&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;py_confirmed_rel&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;py_df&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;loc&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;py_df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;type&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;confirmed&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;groupby&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;country&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;agg&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;cases&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;sum&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;rename&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;columns&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;cases&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;confirmed&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Join population information.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;py_confirmed_rel&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;py_confirmed_rel&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;countries19&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;set_index&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;Location&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;on&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;country&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Calculate relative confirmed cases.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;py_confirmed_rel&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;py_confirmed_rel&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;assign&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;confirmed_rel&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;py_confirmed_rel&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;confirmed&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;py_confirmed_rel&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;PopTotal&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;py_confirmed_rel&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;py_confirmed_rel&lt;/span&gt;&lt;span class="p"&gt;[[&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;confirmed&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;confirmed_rel&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;]]&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;py_confirmed_rel&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sort_values&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;confirmed_rel&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;ascending&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;head&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## confirmed confirmed_rel
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## country
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Andorra 4038 52.342312
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Bahrain 79975 48.730657
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Qatar 130965 46.243544
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Israel 309413 36.318753
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Holy See 27 33.128834
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Panama 128515 30.264174
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Kuwait 120927 28.743710
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Peru 883116 27.164056
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Montenegro 16629 26.479805
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;## Belgium 305409 26.466797
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;Before concluding, I have to admit that I have been using &lt;code&gt;tidyverse&lt;/code&gt; for several years, and I am very used to its syntax. Therefore, it is no wonder that I feel much more comfortable with it than with &lt;code&gt;pandas&lt;/code&gt;&amp;rsquo; . Being said that, I find the latter to be quite straightforward and relatively easy to use and memorise (I will need to check this post and
for reference). However, I have two main concerns about pandas:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Compared to &lt;code&gt;R&lt;/code&gt;, which is more succinct, pandas requires to type many times the data frame&amp;rsquo;s name. This makes it more prone-error and slow to type, but also more difficult to read. If you want to avoid typing it as much, you need to chain a number of methods that result in very long lines, which makes it difficult to comment and it is, again, more difficult to read.&lt;/li&gt;
&lt;li&gt;I find it somewhat overwhelming that there are many ways to perform same task in &lt;code&gt;pandas&lt;/code&gt;. I believe Ted Petrou&amp;rsquo;s advice on learning the
is a good advice, as it makes things simpler.&lt;/li&gt;
&lt;li&gt;I miss the magritte&amp;rsquo;s pipe operator, but I guess I should change my mindset when using python.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Regarding &lt;code&gt;reticulate&lt;/code&gt;: I think it is very promising. While &lt;code&gt;jupyter notebooks&lt;/code&gt; can be used with different kernels such as &lt;code&gt;Julia&lt;/code&gt;, &lt;code&gt;python&lt;/code&gt; or &lt;code&gt;R&lt;/code&gt; (hence its name), there is no way to combine different kernels in the same notebook, at least that I am aware of. This means that only one language per notebook can be used. On the other hand, reticulate allows you to use different languages within the same document (a markdown file, which I prefer it over jupyter notebooks, by the way). Admittedly, I do not know if that is a common scenario, but it has proven to be very useful for a post like this one.&lt;/p&gt;
&lt;p&gt;Being said that, I admit that I expected that I could use one variable from &lt;code&gt;python&lt;/code&gt; and use it in &lt;code&gt;R&lt;/code&gt;, if that makes any sense at all. However, that&amp;rsquo;s not possible, as both languages are isolated, which I assume is the logical way (I assume is not straightforward at all to convert from one &lt;code&gt;python&lt;/code&gt; &lt;code&gt;list&lt;/code&gt; to an &lt;code&gt;R&lt;/code&gt; &lt;code&gt;vector&lt;/code&gt;, for example).&lt;/p&gt;
&lt;div class="footnotes" role="doc-endnotes"&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id="fn:1"&gt;
&lt;p&gt;If there is something that I have fallen in love with &lt;code&gt;python&lt;/code&gt; so far is how convenient and easy are to create virtual environments from a simple &lt;code&gt;yaml&lt;/code&gt; file listing all the dependencies. This is something that would be very useful in &lt;code&gt;R&lt;/code&gt;, too and I should explore in the future: I know &lt;code&gt;packrat&lt;/code&gt; is there for this purpose, but it is not as fast and easy to deal with as &lt;code&gt;conda&lt;/code&gt; environments. On the other hand, if I am not mistaken, &lt;code&gt;conda&lt;/code&gt; also has &lt;code&gt;R&lt;/code&gt; libraries, so it may be possible to create conda environments for R, too.&amp;#160;&lt;a href="#fnref:1" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</description></item></channel></rss>