NA446A Python for data you can trust

Sample exam, Python part

Ten questions in the format of the real exam. Answer without AI and without running code, then check.

Work on paper, as in the exam hall: read the code and predict. Give yourself 15 minutes. The real exam has more questions and also covers Part I.

1. What does the last line print?

import pandas as pd
df = pd.DataFrame({"id": [1, 2, 3], "wage": [100, 120, 90]})
print(df.shape)

A. (2, 3) B. (3, 2) C. 6 D. (3,)

2. Why does the course use paths like ROOT / "data" / "raw" instead of C:/Users/anna/Desktop/project/data/raw?

A. pandas only accepts relative paths B. The project then runs on any computer where the folder is copied C. Windows cannot read long paths D. It makes the code run faster

3. What does the last line print?

years = [2019, 2020, 2021, 2022]
print(years[1])

A. 2019 B. 2021 C. An error D. 2020

4. A CSV column holds values like 12, 15 and “****” (a missing-value code). What type does pandas give the column when it reads the file?

import io
import pandas as pd
df = pd.read_csv(io.StringIO("id,income\n1,12\n2,15\n3,****\n"))
kind = "text" if df["income"].dtype.kind in "OUT" else "number"

A. category B. text C. number D. date

5. How many rows does high have?

import pandas as pd
df = pd.DataFrame({"mpg": [30.0, None, 18.0, 26.0]})
high = df[df["mpg"] > 25]

A. 2 B. 4 C. 3 D. 1

6. How many columns does wide have?

import pandas as pd
long = pd.DataFrame({"pid": [1, 1, 2, 2], "year": [2023, 2024, 2023, 2024], "wage": [100, 110, 90, 95]})
wide = long.pivot(index="pid", columns="year", values="wage")

A. 1 B. 3 C. 4 D. 2

7. How many rows does long have?

import pandas as pd
wide = pd.DataFrame({"pid": [1, 2, 3], "wage2023": [100, 110, 90], "wage2024": [105, 112, 95]})
long = wide.melt(id_vars="pid", var_name="year", value_name="wage")

A. 3 B. 2 C. 6 D. 9

8. Why does the course’s run_all log print the number of rows in and out at every step?

A. Because pandas requires it B. To save memory C. So that an unexpected drop or duplication of rows is seen at the step where it happens D. To make the script run faster

9. Gentzkow and Shapiro (2014) describe a precise test of whether a project’s output is replicable. What is it?

A. Ask a coauthor to read the code B. Save every file with the date and your initials C. Run the regressions twice and compare D. Delete all output files and reproduce them by running the one script that runs the directory

10. Quarterly GDP is 100 in 2023Q1 and 103 in 2024Q1. Which line gives year-on-year growth in per cent for quarterly data?

A. 100 * gdp.pct_change(4) B. gdp.pct_change(12) C. 100 * gdp.diff(4) D. 100 * gdp.pct_change(1)

Answers and why

Show the answers (try first)
  1. B. shape is (rows, columns): three rows, two columns.
  2. B. An absolute path works on one computer only; a path from the project root travels with the project.
  3. D. Python counts from 0, so position 1 is the second element (MATLAB would count from 1).
  4. B. One non-numeric entry makes the whole column text; replace the code with missing and convert with pd.to_numeric.
  5. A. A missing value compared with > gives False, so the missing row is dropped (Stata would treat . as larger than any number and keep it).
  6. D. Long to wide gives one column per year, with pid as the index (Stata: reshape wide).
  7. C. Wide to long gives one row per person and year: 3 x 2 = 6 (Stata: reshape long).
  8. C. Row counts are the cheapest check there is: a merge that explodes or a filter that drops too much shows up immediately.
  9. D. In the Automation chapter, deleting the outputs and rebuilding them from the run script is ‘the precise sense in which the output is now replicable’; dates and initials are what version control replaces.
  10. A. Year on year compares with the same quarter a year earlier: four quarters back.