Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 7 additions & 1 deletion CATALOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,19 +6,25 @@

The dataset registry, **auto-generated** from the sidecar manifests (`lectures/*.yml`). Do not edit by hand — run `python scripts/build_catalog.py`. A dataset appears here once it has a manifest, which may be before its consuming lectures are repointed — an empty **Used by** column means the file is here and documented but no lecture reads it from this repo yet. Files still to migrate are tracked in [PLAN.md](PLAN.md).

**18 datasets** · 18 read by lectures today · 5.3 MB total · 17 permitted / 1 restricted redistribution
**24 datasets** · 18 read by lectures today, 6 awaiting repoint · 110.0 MB total · 19 permitted / 5 restricted redistribution

| Dataset | Class | Source | Licence | Redist. | Integrity | Builder | Size | Used by |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [**SCF_plus_mini.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/SCF_plus_mini.csv)<br><sub>SCF+ mini — net wealth, income and survey weights, 1950-2016</sub> | constructed | [SCF+ (Kuhn, Schularick and Steins) — an extension of the Survey of Consumer Finances](https://www.journals.uchicago.edu/doi/10.1086/708815) | | ✅ permitted | ⚠️ unverifiable | committed-frozen | 31.3 MB | — |
| [**SCF_plus_mini_no_weights.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/SCF_plus_mini_no_weights.csv)<br><sub>SCF+ mini, weight-expanded — net wealth and income, 1950-2016</sub> | constructed | [SCF+ (Kuhn, Schularick and Steins) — an extension of the Survey of Consumer Finances](https://www.journals.uchicago.edu/doi/10.1086/708815) | | ✅ permitted | ⚠️ unverifiable | committed-frozen | 72.4 MB | — |
| [**ames_house_prices.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/ames_house_prices.csv)<br><sub>Ames, Iowa — residential house sales, 2006-2010</sub> | constructed | [Ames Housing data (De Cock 2011), Journal of Statistics Education](http://jse.amstat.org/v19n3/decock.pdf) | | ✅ permitted | ✅ verified | ✅ committed | 75.2 KB | [lecture-python-intro · observed_distributions.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/observed_distributions.md)<br>[lecture-python-intro · fitting_distributions.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/fitting_distributions.md) |
| [**assignat.xlsx**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/assignat.xlsx)<br><sub>French Revolution — assignat issues, budgets and seigniorage (Sargent-Velde)</sub> | verbatim | [Sargent and Velde, "Macroeconomic Features of the French Revolution" — supporting spreadsheets](https://www.journals.uchicago.edu/doi/10.1086/261992) | | ✅ permitted | ⚠️ unverifiable | n/a (verbatim) | 204.6 KB | [lecture-python-intro · french_rev.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/french_rev.md)<br>[lecture-wasm · french_rev.md](https://github.com/QuantEcon/lecture-wasm/blob/main/lectures/french_rev.md) |
| [**caron.npy**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/caron.npy)<br><sub>French Revolution — monthly specie value of the assignat, 1791-1796</sub> | constructed | unrecorded | | ✅ permitted | ⚠️ unverifiable | ⚠️ unrecovered | 1.1 KB | [lecture-python-intro · french_rev.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/french_rev.md)<br>[lecture-wasm · french_rev.md](https://github.com/QuantEcon/lecture-wasm/blob/main/lectures/french_rev.md) |
| [**chapter_3.xlsx**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/chapter_3.xlsx)<br><sub>The Ends of Four Big Inflations — appendix tables, transcribed</sub> | constructed | [Sargent, "Rational Expectations and Inflation", chapter 3 appendix tables](https://press.princeton.edu/books/paperback/9780691158709/rational-expectations-and-inflation) | | ✅ permitted | ⚠️ unverifiable | ⚠️ unrecovered | 71.6 KB | [lecture-python-intro · inflation_history.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/inflation_history.md)<br>[lecture-wasm · inflation_history.md](https://github.com/QuantEcon/lecture-wasm/blob/main/lectures/inflation_history.md) |
| [**cities_brazil.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/cities_brazil.csv)<br><sub>World Population Review — Brazilian city populations, 2023</sub> | verbatim | [World Population Review — cities in Brazil](https://worldpopulationreview.com/countries/cities/brazil) | | ⚠️ restricted | ⚠️ unverifiable | n/a (verbatim) | 17.5 KB | — |
| [**cities_us.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/cities_us.csv)<br><sub>World Population Review — US city populations, 2023</sub> | verbatim | [World Population Review — US cities](https://worldpopulationreview.com/us-cities) | | ⚠️ restricted | ⚠️ unverifiable | n/a (verbatim) | 47.0 KB | — |
| [**countries.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/countries.csv)<br><sub>WorldData.info country reference table</sub> | verbatim | [WorldData.info — country data downloads](https://www.worlddata.info/downloads/) | Proprietary — © WorldData.info, all rights reserved | ⚠️ restricted | ⚠️ unverifiable | n/a (verbatim) | 48.4 KB | [lecture-python-programming · pandas_panel.md](https://github.com/QuantEcon/lecture-python-programming/blob/main/lectures/pandas_panel.md)<br>[lecture-python.myst · pandas_panel.md](https://github.com/QuantEcon/lecture-python.myst/blob/main/lectures/pandas_panel.md) |
| [**dette.xlsx**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/dette.xlsx)<br><sub>French Revolution — public debt, military spending and revenues (Sargent-Velde)</sub> | verbatim | [Sargent and Velde, "Macroeconomic Features of the French Revolution" — supporting spreadsheets](https://www.journals.uchicago.edu/doi/10.1086/261992) | | ✅ permitted | ⚠️ unverifiable | n/a (verbatim) | 617.2 KB | [lecture-python-intro · french_rev.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/french_rev.md)<br>[lecture-wasm · french_rev.md](https://github.com/QuantEcon/lecture-wasm/blob/main/lectures/french_rev.md) |
| [**employ.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/employ.csv)<br><sub>Eurostat employment in Europe — by age and sex, 2007–2016</sub> | constructed | [Eurostat — Employment database](https://ec.europa.eu/eurostat/data/database) | Eurostat reuse (Commission Decision 2011/833/EU) | ✅ permitted | ⚠️ unverifiable | ⚠️ unrecovered | 1.6 MB | [lecture-python-programming · pandas_panel.md](https://github.com/QuantEcon/lecture-python-programming/blob/main/lectures/pandas_panel.md)<br>[lecture-python.myst · pandas_panel.md](https://github.com/QuantEcon/lecture-python.myst/blob/main/lectures/pandas_panel.md) |
| [**epl_match_goals.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/epl_match_goals.csv)<br><sub>English Premier League — full-time scores, 2015-16 to 2024-25</sub> | constructed | [openfootball / football.json](https://github.com/openfootball/football.json) | Public domain | ✅ permitted | ✅ verified | ✅ committed | 203.2 KB | [lecture-python-intro · fitting_distributions.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/fitting_distributions.md) |
| [**fig_3.xlsx**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/fig_3.xlsx)<br><sub>French Revolution — figure 3 series (Sargent-Velde)</sub> | verbatim | [Sargent and Velde, "Macroeconomic Features of the French Revolution" — supporting spreadsheets](https://www.journals.uchicago.edu/doi/10.1086/261992) | | ✅ permitted | ⚠️ unverifiable | n/a (verbatim) | 9.2 KB | [lecture-python-intro · french_rev.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/french_rev.md)<br>[lecture-wasm · french_rev.md](https://github.com/QuantEcon/lecture-wasm/blob/main/lectures/french_rev.md) |
| [**forbes-billionaires.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/forbes-billionaires.csv)<br><sub>Forbes Billionaires — individual net worth</sub> | constructed | [Forbes Billionaires](https://www.forbes.com/billionaires/) | | ⚠️ restricted | ⚠️ unverifiable | committed-frozen | 775.7 KB | — |
| [**forbes-global2000.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/forbes-global2000.csv)<br><sub>Forbes Global 2000 — firm size measures</sub> | constructed | [Forbes Global 2000](https://www.forbes.com/lists/global2000/) | | ⚠️ restricted | ⚠️ unverifiable | committed-frozen | 115.6 KB | — |
| [**japan_deaths_by_age.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/japan_deaths_by_age.csv)<br><sub>Japan — deaths by single year of age, 2023</sub> | constructed | [United Nations, Department of Economic and Social Affairs, Population Division — World Population Prospects 2024](https://population.un.org/wpp/downloads) | CC BY 3.0 IGO | ✅ permitted | ✅ verified | ✅ committed | 1.7 KB | [lecture-python-intro · observed_distributions.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/observed_distributions.md)<br>[lecture-python-intro · fitting_distributions.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/fitting_distributions.md) |
| [**japan_earthquakes.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/japan_earthquakes.csv)<br><sub>Japan region — earthquakes of magnitude 5 and above, 2000-2024</sub> | constructed | [Advanced National Seismic System (ANSS) Comprehensive Earthquake Catalog (ComCat), US Geological Survey](https://earthquake.usgs.gov/earthquakes/search/) | US Government work — public domain | ✅ permitted | ✅ verified | ✅ committed | 172.8 KB | [lecture-python-intro · fitting_distributions.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/fitting_distributions.md) |
| [**japan_population_by_age.csv**](https://github.com/QuantEcon/data-lectures/raw/main/lectures/japan_population_by_age.csv)<br><sub>Japan — population by single year of age, 2024</sub> | constructed | [Population Estimates, Statistics Bureau of Japan, Ministry of Internal Affairs and Communications](https://www.stat.go.jp/english/data/jinsui/index.html) | Japan Statistics Bureau terms of use | ✅ permitted | ✅ verified | ✅ committed | 1.3 KB | [lecture-python-intro · prob_dist.md](https://github.com/QuantEcon/lecture-python-intro/blob/main/lectures/prob_dist.md) |
Expand Down
81 changes: 81 additions & 0 deletions builders/generating_mini.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,81 @@
---
jupytext:
text_representation:
extension: .md
format_name: myst
format_version: 0.13
jupytext_version: 1.14.1
kernelspec:
display_name: Python 3 (ipykernel)
language: python
name: python3
---

Regarding converting between ``.ipynb`` and ``.md`` please refer to https://manual.quantecon.org/writing/converting.html
Comment thread
mmcky marked this conversation as resolved.

```{code-cell} ipython3
import pandas as pd
```

```{code-cell} ipython3
var_list = ['yearmerge', # 3-year window
'ffanw', # net wealth (ffafin + ffanfin - tdebt)
'tinc', # total household income, excluding capital gains
'incws', # income from wages, salaries and self-employment
'wgtI95W95', # survey weight
'ffanwgroups', # wealth groups
'tincgroups'] # income groups
```

Rename the variables needed.

```{code-cell} ipython3
var_names_new = 'year', 'n_wealth', 't_income', 'l_income', 'weights', 'nw_groups', 'ti_groups'
```

```{code-cell} ipython3
df = pd.read_stata('https://github.com/QuantEcon/high_dim_data/blob/main/SCF_plus/SCF_plus.dta?raw=true')
```

```{code-cell} ipython3
df = df[[*var_list]]
df1=df.astype({'yearmerge': int}).dropna()
df1.columns = var_names_new
```

```{code-cell} ipython3
df1
```

Export the dataset with weights.

```{code-cell} ipython3
# df1.to_csv('SCF_plus_mini.csv', index=None) # use it when you want to export the weighted data
```

Generate and export the dataset without weights.

```{code-cell} ipython3
counts = list(round(df1['weights']))
df1["weights"] = counts
```

```{code-cell} ipython3
df2 = df1.loc[df1.index.repeat(df1.weights)].reset_index(drop=True)
```

```{code-cell} ipython3
df2 = df2.drop(columns=['weights'])
```

```{code-cell} ipython3
df2
```

```{code-cell} ipython3
# df2.to_csv('SCF_plus_mini_no_weights.csv', index=None) # use it when you want to export the non weighted data
```

```{code-cell} ipython3

```
156 changes: 156 additions & 0 deletions builders/webscrape_forbes.ipynb
Original file line number Diff line number Diff line change
@@ -0,0 +1,156 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# parse_forbeslists\n",
"\n",
"This notebook \n",
"- parses Forbes richest lists and Forbes global 2000 list and\n",
"- saves them as csv files."
]
},
{
"cell_type": "code",
"execution_count": 1,
"metadata": {},
"outputs": [],
"source": [
"import requests\n",
"import pandas as pd\n",
"from pathlib import Path\n",
"from pandas import DataFrame"
]
},
{
"cell_type": "code",
"execution_count": 2,
"metadata": {},
"outputs": [],
"source": [
"# Forbes lists\n",
"lists = [ \n",
" { 'type': 'person', 'year': 2020, 'uri': 'billionaires' }, # World richest\n",
" # { 'type': 'person', 'year': 2020, 'uri': 'forbes-400' }, # American richest 400\n",
" # { 'type': 'person', 'year': 2020, 'uri': 'hong-kong-billionaires' }, # Hong Kong richest 50\n",
" # { 'type': 'person', 'year': 2020, 'uri': 'australia-billionaires' }, # Australia richest 50\n",
" # { 'type': 'person', 'year': 2020, 'uri': 'china-billionaires' }, # China richest 400\n",
" # { 'type': 'person', 'year': 2020, 'uri': 'taiwan-billionaires' }, # Taiwan richest 50\n",
" # { 'type': 'person', 'year': 2020, 'uri': 'india-billionaires' }, # India richest 100\n",
" # { 'type': 'person', 'year': 2020, 'uri': 'japan-billionaires' }, # Japan richest 50\n",
" # { 'type': 'person', 'year': 2020, 'uri': 'africa-billionaires' }, # Africa richest 50\n",
" # { 'type': 'person', 'year': 2020, 'uri': 'korea-billionaires' }, # Korea richest 50\n",
" # { 'type': 'person', 'year': 2020, 'uri': 'malaysia-billionaires' }, # Malaysia richest 50\n",
" # { 'type': 'person', 'year': 2020, 'uri': 'philippines-billionaires' }, # Philippines richest 50\n",
" # { 'type': 'person', 'year': 2020, 'uri': 'singapore-billionaires' }, # Singapore richest 50\n",
" # { 'type': 'person', 'year': 2020, 'uri': 'indonesia-billionaires' }, # Indonesia richest 50\n",
" # { 'type': 'person', 'year': 2020, 'uri': 'thailand-billionaires' }, # Thailand richest 50\n",
" # { 'type': 'person', 'year': 2020, 'uri': 'self-made-women' }, # American richest self-made women\n",
" # { 'type': 'person', 'year': 2018, 'uri': 'richest-in-tech' }, # tech richest\n",
" # { 'type': 'person', 'year': 2020, 'uri': 'hedge-fund-managers' }, # hedge fund highest-earning\n",
" # { 'type': 'person', 'year': 2020, 'uri': 'powerful-people' }, # world powerful\n",
" # { 'type': 'person', 'year': 2020, 'uri': 'power-women' }, # world powerful women\n",
" # { 'type': 'person', 'year': 0, 'uri': 'rtb' }, # real-time world billionaires\n",
" # { 'type': 'person', 'year': 0, 'uri': 'rtrl' }, # real-time American richest 400\n",
"]\n",
"\n",
"url = 'http://www.forbes.com/ajax/list/data'\n",
"SOURCES_DIR = Path('./sources')\n",
"\n",
"for forbes_list in lists:\n",
" response = requests.get(url, params=forbes_list)\n",
Comment thread
mmcky marked this conversation as resolved.
"\n",
" if not SOURCES_DIR.exists():\n",
" SOURCES_DIR.mkdir(exist_ok=True, parents=True)\n",
"\n",
" DataFrame(response.json()).to_csv('forbes-{}.csv'.format(forbes_list['uri']))"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Then Forbes Global 2000 for the largest 2000 firms globally."
]
},
{
"cell_type": "code",
"execution_count": 3,
"metadata": {},
"outputs": [],
"source": [
"headers = {\n",
" \"accept\": \"application/json, text/plain, */*\",\n",
" \"referer\": \"https://www.forbes.com/global2000/\",\n",
" \"user-agent\": \"Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/87.0.4280.67 Safari/537.36\",\n",
"}\n",
"\n",
"cookies = {\n",
" \"notice_behavior\": \"expressed,eu\",\n",
" \"notice_gdpr_prefs\": \"0,1,2:1a8b5228dd7ff0717196863a5d28ce6c\",\n",
"}\n",
"\n",
"api_url = \"https://www.forbes.com/forbesapi/org/global2000/2020/position/true.json?limit=2000\"\n",
"response = requests.get(api_url, headers=headers, cookies=cookies).json()\n",
"\n",
"sample_table = [\n",
" [\n",
" item[\"organizationName\"],\n",
" item[\"country\"],\n",
" item[\"revenue\"],\n",
" item[\"profits\"],\n",
" item[\"assets\"],\n",
" item[\"marketValue\"]\n",
" ] for item in\n",
" sorted(response[\"organizationList\"][\"organizationsLists\"], key=lambda k: k[\"position\"])\n",
"]\n",
"\n",
"dfff = pd.DataFrame(sample_table, columns=[\"Company\", \"Country\", \"Sales\", \"Profits\", \"Assets\", \"Market Value\"])\n",
"dfff.to_csv('forbes-global2000.csv')"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": []
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": []
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": []
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3",
"version": "3.10.9"
}
},
"nbformat": 4,
"nbformat_minor": 5
}
Loading
Loading