NA446A Python for data you can trust

Readings

Short and open. Read the parts named here; each takes 15–25 minutes.

The exam includes one or two questions on these readings.

When Read Why
Before session 1 (required) Bryan, K. A. (2025). Social Science PhD Tech Stack. Memo, University of Toronto. Read the introduction and the three “Reproducibility” parts: Starting a Project, Version Control (skim the git commands) and What to Use for Code. What reproducible, efficient and open mean, and how a leading economist sets up every project
Before session 1 (optional) Bastani, H. et al. (2025). Generative AI without guardrails can harm learning. PNAS 122(26): e2422633122 (open version). The abstract only. Why the course’s AI rules look the way they do
Before session 2 Turrell, A. et al. Coding for Economists, Data Analysis Quickstart. Read from Loading data to Add transformed columns; skip Randomly selecting a sample and Rename. The pandas operations of session 2, explained for economists
Before session 3 Gentzkow, M. and J. M. Shapiro (2014). Code and Data for the Social Sciences: A Practitioner’s Guide. Chapter 5, Keys, pp. 18–21. Why merges go wrong, and the rule that prevents it: unique, non-missing keys
Before session 4 Gentzkow and Shapiro (2014), Chapter 2, Automation, pp. 6–10. One script that runs everything, and the delete-the-outputs test

On Bryan’s memo: he recommends Cursor as the editor; we use VS Code, which Cursor is built on, so everything carries over. He is also dismissive of Stata and MATLAB. Read that as one economist’s view: MATLAB is standard for numerical work in finance and macro (Part I), and Stata is still common in applied research and government. The skills in this course carry across all of them.

Reference, not reading: the pandas user guide pages on missing data, merging and reshaping. Look things up there when you need them.