Data Dojo Würzburg 11
DataDojo@Lunch
April 2022
- When: Thursday, April 14th, 2022 at 12:00pm
- Where: Zoom
- Zoom:
- Link
- Meeting ID: 681 8517 5809
- Password: 853667
- Info: DataDojo Website, Repo
Participants
Please add your name to the list (click the pen icon at the top left to edit) if you plan to come. And please remove it if you can not make it. Feel free to add your preferred tool or programming language.
- Markus (R or julia)
- Robin R (R/julia)
- Andi (Julia)
Dataset
Annual German weather data. Compiled from the Climate Data Center. We have a table like this:
station_id | year | temperature | precipitation | sunshine | station_name | lat | long | altitude |
---|---|---|---|---|---|---|---|---|
5703 | 2020 | 9.7 | 527.3 | 5.1 | Würzburg | 49.8 | 9.96 | 268 |
… |
for more than 5000 weather stations across Germany and a timespan of over 100 years (for some stations and measurements).
Specific task for today
30 Day Chart Challenge Special
Create a chart in the category “Relationships - 3-dimensional”
The stations are already 3D (latitude, longitude, altitude). But we want to plot some relationships of the 3D data regarding temperature, precipitation and sunshine duration (maybe over time).
Task: create a chart with the provided data set, that matches the category: Relationships, 3-dimensional.
Question Pool:
- Generic
- What kind of information is stored in the table(s)?
- How much data is missing?
- Is the dataset clean or are there any clear outliers?
- How can the different datasets be combined?
- How to visualize the results in a suitable way?
- Specific
- Add your own questions
- Further Ideas
- Add your own ideas
Collaborative Tools and Workflow
For Notebooks (R, python, julia, js, …) with real time collaboration CoCalc seems to be the best option right now. It worked great the last couple of times so we’ll stick to it for now. You need to register an account there (it is free).
Future Suggestions
Add your suggestions to the list and :+1: to the end of a line you are interested in
Data Sets
- Results of the Bundestagswahl 2021
- Weather data throughout Germany over time (incl. temperature, precipitation, …): https://www.dwd.de/DE/leistungen/cdc_portal/cdc_portal.html
- German Mikrozensus
- Kaggle Titanic or Tabular Playground or Meta Kaggle
- World Trade Data (Open Trade Statistics)
- Open Citation Data
- Top 100 charts + Audio Features
- Emoji Usage :hugging_face::heart::laughing:
Tools/Languages
Skills
- interactive maps
- dashboards
- animations
Data Sources
all data types are welcome, including tables, images, videos, sounds, DNA, …
- TidyTuesday
- Our World in Data (R package: owidR), Sustainable Development Goals
- Open Data Initiatives (Würzburg, Germany, Statistisches Bundesamt, Europe, APIs)
- Data is plural
- Awesome Public Datasets
- Kaggle Datasets or Competitions, e.g. SLICED
- tsibbledata: Time Series Datasets
- R-text-data: Text Datasets, ready to use in R
- data.world
- Statista - the University of Würzburg has a campus license
- Open Legal Data
- Bundestag Data (e.g. poll results, deputies, wahl-o-mat, inspirational blog post)
- Deutsche Digitale Bibliothek (API, old newspapers from Germany)
- Earth Observation: Satellite Image Time Series
- Machine Learning Datasets
- Internation (Student) Assessment Data (TIMSS, PIRLS, PISA, …)
- (Medical) Imaging Datasets, MedMNIST
- Inspirational Notebooks on Observable
- Ski resort statistics :skier: