Gov 50
  • Syllabus
  • Schedule
  • Staff
  • Office Hours & Study Halls
  • Assignments
  • Resources
  • Canvas
  • Ed

On this page

  • Course objectives
    • A note on the structure and difficulty of this course
    • Expectations and time commitment
    • Prerequisites
    • Credit
  • Course structure
    • Textbooks & readings
    • Warmup questions
    • In-Person participation
    • Classes
    • Section
    • Study Halls
    • Homework
    • Section replication project
    • Final exam
    • Grades
  • Course policies
    • Laptops and written materials
    • Academic honesty
    • Coding and AI
    • Late policy
    • Regrading policy
    • Office hours, study halls, and availability
  • Course materials
    • Books
    • Computing
  • Accessibility
  • Acknowledgements

Syllabus

Instructor

  • Dr. Sam Fuller
  • CGIS Knafel 428
  • sfuller@fas.harvard.edu
  • Schedule an appointment

Teaching fellows

  • (Head TF) Nicholas Conroy
  • Antonio Camara
  • Amin Karimi
  • Shashwat Kulkarni
  • Zhiyu Li
  • Tianyu Qiao
  • George Yean

Course details

  • Mon/Wed
  • September 2nd-December 18th, 2026
  • 3:00–4:15 PM
  • Emerson 105
  • Ed Discussion Board

Course assistants

  • Avi Agarwal
  • Frida Bravo Lopez
  • Daniel Cabrera
  • Alan Cai
  • Abril Diaz
  • Alex Draghia
  • Shuxin Ho
  • Harrison Huang
  • Jenna Jiang
  • Mateo Quintero
  • Megan Martinez
  • Lenny Pische
  • Amy Tan
  • Jahir Tineo Castillo
  • Reed Trimble
  • Todd Zhou

Course objectives

In this course you will learn the basics of quantitative analysis (data science) in the social sciences. These basics can be broken down into two broad categories: 1) understanding the theoretical basis of data analysis and how that influences our analytical decisions; and 2) learning how to actually collect, manage, and analyze data (as well as report the results of said analyses).

At the end of the course you should be able to read and digest most social science articles by learning to:

  • Interpret common analyses in the social sciences (e.g., regression tables)
  • Evaluate claims about causality (correlation ≠ causation)
  • Identify and understand common causal designs and interrogate their validity
  • Understand what uncertainty means in data analysis

Additionally, by taking this knowledge and applying it you will be able to:

  • Use common tools for data analysis, mainly R and RStudio
  • Wrangle messy data into tidy, usable forms
  • Summarize and visualize data in a compelling way
  • Take a research question from inception to analysis
  • Use correlations and linear regressions to analyze data
  • (we may also get to implementing causal designs…)

All of this together will hopefully generate an interest (and maybe passion!) for data analysis through creating a welcoming community among students and instructors.

A note on the structure and difficulty of this course

Learning statistical/social-science/political methodology is difficult, there’s no two ways about it. You will be tasked with essentially learning two foreign languages: statistics/data-analysis and R. Continuing with this analogy of language, the goal here is to learn how to have and understand conversations, not to simply learn basic turns of phrase and rote grammatical rules.

That being said, each and every one of you is capable of grasping this material, both conceptually and computationally. If you do really engage with this course then you will come away with a great deal of knowledge that will prepare you to interpret, critique, and build upon social science research.

To that end, this course will look quite different from most other lecture-based courses (though I obviously can’t speak for other disciplines, I’m sure they’re doing cool stuff too, at least some of them…). This design can be described as “active learning” or a “flipped classroom” or “collaborative learning” (pick your poison, I personally like “active”).

The goal of active learning is to, first and foremost, improve your learning, but it does require buy-in from you, the student. To truly learn the content here—and trust me, you will if you put in the effort—you must show up to class, section, and study-hall prepared and actively (ha) choose to engage (details about this are available below). Put simply, class-time (and sections and study-halls) will be focused on much more than simply covering material.

I’m very excited to be teaching you all (along with my wonderful TFs and CAs) and I think you all will learn a lot and enjoy it too!

Expectations and time commitment

In this course, you will be expected to:

  • complete readings and answer warmup questions for each class,
  • participate in in-class discussions,
  • attend and participate in weekly sections,
  • complete six in-section coding problem sets (every other week),
  • complete 23 homework assignments,
  • take one in-person, closed-book exam, and
  • contribute to a section-wide final research project.

The course is designed to require 10 hours of work per week, roughly divided as:

  • 2 hours reading
  • 0.5 hours on pre-class warmup questions
  • 2.5 hours in the classroom
  • 1 hour in section
  • 4 hours on homework.

For reading week before your exam you are expected to study for another 10 hours.

Prerequisites

We will assume you are coming in (relatively) blind given that the course has no prerequisites. If you don’t have much experience with downloading, installing, and interacting with software on your computer, you should probably spend additional time on those aspects (particularly early on in the course) to make sure everything proceeds smoothly. We will be providing an Intro to R handout and will dedicate the first meeting of sections to getting you up to speed with R.

Credit

This course satisfies the Methods requirement for the Government Department, the Quantitative Reasoning with Data requirement in the Harvard College Curriculum, and also counts toward the Government Department’s Data Science track.

Course structure

Textbooks & readings

Back when I was in undergrad I was (mostly) able to get my textbooks through Inter-Library Loan, mooching, or other means. The few times that I was forced to spend $100 (or more) on textbooks I was quite annoyed…

Consequently, I have assigned two completely free textbooks for this course: Regression and Other Stories by Andrew Gelman (take a look at his blog for really great thinking on statistics and science), Jennifer Hill, & Aki Vehtari as well as The Effect: An Introduction to Research Design and Causality by Nick Huntington-Klein.1 You are more than welcome to purchase hard copies of either book but you are by no means required (and technically the online copies are the most up-to-date).

1 Special thanks to Jack Rametta and Chris Hare for their book recommendations here.

You will be expected to have completed the assigned readings by the time you attend class. Relatedly, warmup assignments will be due 1 hour before each class and will correspond directly to the readings assigned for each class.

Warmup questions

You will be required to answer a small set of questions for every class meeting (due 1 hour before class). These will be graded on completion not on accuracy. That being said, if you provide answers that fail to signal that you’re engaging with the material (e.g., a response unrelated to the material) you will not be given credit.

In-Person participation

Your participation in class, section, and study-halls is not only a requirement for the course itself (your grade depends on it) but also for you to learn the material. In-person participation will be evaluated by your TF and CAs.

Missing class/section or failing to participate will hurt your grade in the class, each unexcused absence will reduce your overall grade by 0.5pts and failing to participate significantly will decrease your grade by 0.25pts.

  1. In-class participation will be monitored by TFs and CAs
  2. In-section participation will be monitored by TFs
  3. Study hall participation will be monitored by CAs
    • Note: Attendance is not mandatory for Study Halls, but can be used to supplement other participation grades.
NoteExample grade with participation policy

So, for example, if you have earned a 92 in the class, but you miss 2 classes and 2 sections (4 unexcused absences) and fail to participate significantly 4 times, then your final grade would actually be an 89 (92 - ((4\times0.5)+(4\times0.25))=92 - (2+1)=89).

Classes

ImportantWhat to bring to class

As I mentioned above, for every class you should bring a) your laptop, b) a notebook (broadly construed), and c) questions about the materials we’ve covered.

We will meet twice a week for regular classes. As I mentioned above, these classes will not follow standard lecture structures but will rather be focused on active learning. Most classes will be comprised of:

  1. Story
  2. Class-participation activity
  3. Class discussion of questions relating to readings and homework
  4. Code demonstration
  5. Drill (do something yourself)
  6. Small-group discussion problems

Classes will alternate between focusing on statistics and research design (including causal inference), though this distinction will likely blur as we continue through the course. The class schedule is available here.

NoteSeating

On the first day of class we will have the TFs and CAs separated into different sections of the lecture hall (they will have a piece of paper with their names) and you will be required to sit with your assigned TF (the one you’re taking section with).

You will be required to sit in these areas for the whole semester so that we are able to a) collect homework quickly and efficiently and (more importantly) b) facilitate class activities and group discussions.

Section

Sections will be a mandatory part of the course and attendance will be taken. Students will be assigned to their section randomly based on availability. Sections will alternate weekly between: 1) covering material and developing a section-level project; and 2) hosting in-person problem sets.

For material and project weeks, sections will be focused on helping you understand the material we’ve been covering in class and in developing a research project that everyone in your section will contribute to.

NoteIn-section coding problems

For in-person coding problem set weeks, sections will be focused on increasing your skills at using R and interpreting related code. You will be encouraged to work with other students in section and the TF will be helping the class move through the assignment, checking in throughout and answering questions. These assignments will be graded similar to the homework rubric (for each question: 0 not attempted, 0.5 attempted but not thoroughly, 1 thoroughly attempted).

Additionally, TFs will not solve the problem sets for you. During section we recommend you, in order, 1) attempt on your own, 2) ask for help from your partner (or attempt with them), and then 3) ask for help from the TF or CA.

Study Halls

Study Halls are a mix of office hours and drop-in tutoring sessions. CAs will hold a table—usually at one of the house dining halls or common rooms—and will answer questions relating to assignments and course material. We highly recommend regularly attending a Study Hall to complete assignments along with your peer study group, attempting to answer questions on your own and looking to a CA when you are stuck.

You can find a list of the Study Hall times on the Study Hall Schedule page.

ImportantCAs will not do the homework for you

CAs are instructed to help guide you through homework assignments, not to show you how to answer a question from start to finish. You must actually do the work to complete the assignments.

Homework

To quote Matt Blackwell, a previous instructor of this course, “Only reading about data science is about as instructive as reading a lot about hammers or watching someone else wield a hammer. You need to get your hands on a hammer or two.” In that vein, you will have homework assignments due at the start of each class period (printed and handed in to your TFs and CAs).

ImportantHomework is due in-person, no exceptions

Homework must be turned in in-person by the start of class, there will not be a way to turn in homework online. There will be no exceptions made to this policy, except for excused absences with prior notice.

Some will be focused on concepts, while others will focus on coding, and others still will be a mix of both. Homework will be posted no later than a week before it’s do (but may be posted earlier).

We encourage you to join peer study groups (we will attempt to facilitate this within section with TFs and CAs) and use them to work on homework and study for the exam. However, the work that you submit should be your own.

NoteGrading policy

Each homework question will be graded as either a 0 (not attempted), 0.5 (attempted but incorrect or incomplete), or 1 (correct); no additional partial credit will be given.

We will drop the two lowest homework grades automatically.

Section replication project

Unlike previous iterations of this course, there will not be an individual final project, but rather a section-level replication (research) project.

For this project, your section will divide into 3-4 research groups, headed by your TF, and tasked with exploring and extending a published paper. You, your classmates, your TF, and your CA will decide upon a paper to replicate in the first few meetings of section and will spend the rest of the semester conducting the replication and an extension (an additional research question inspired by the paper). This process culminates into a completed project that will be presented to a different TF (or myself) in the last section meeting.

50% of your grade will depend on your section participation and contribution to the project and the other 50% will be determined by the presentation.

Your section will be divided into three groups, each tasked with a different part of the overall research question. Each group will be responsible for, for both the replication and extension:

  1. Data collection and cleanup
  2. Statistical summaries and analyses
  3. Visualizations of data and analyses
  4. Interpretation of results and figures
  5. Linking the results to the broader question in your section
  6. Clearly defining the scope of your results (limitations, how your analyses could be improved)

Project presentations should be ~45 minutes long, with the first ten minutes outlining the combined replication results from the different groups, and the following thirty-five minutes divided between the different extensions conducted by each research group. After the presentations, students will field questions from the attending TF (and/or myself).

TipThis process will be highly structured (and fun!)

While this may all sound like a lot, you will be collaborating with members of your section and both your TF and CA will guide you through this process. This will be fun! You’re going to start producing knowledge rather than just consuming it (that’s really cool!).

Milestone Due Date2
Paper for replication selected Section Meeting 2
Paper outlined, replication plan drafted Section Meeting 4
Data cleaned, first summaries of data Section Meeting 6
Main analyses completed, draft of visualizations Section Meeting 8
Replication draft and extension proposal due for peer review Section Meeting 9
Peer review of other groups Section Meeting 10
Incorporate feedback, finalize presentation Section Meeting 11
Presentation to other TF (or Dr. Fuller) Section Meeting 12

2 These meetings correspond (roughly) to the modules of the course. So Section Meeting 2 will correspond to Module 2, Meeting 3 module 3, and so on. Check out the schedule for more info.

ImportantA reminder about AI and academic honesty

Remember, AI is expressly prohibited for any writing or slide generation. Anything written for this project should only be written by you and/or in collaboration with other students, your TF, or your CA. Again, coding can be done by AI, but we highly recommend against doing this for your own sake.

Final exam

There will be one three-hour, in-person exam held during exam week. We will provide a practice exam before the start of reading week and Study Halls will run throughout that week. The exam will test you on:

  1. your understanding of how to use statistics for quantitative social science
  2. your understanding of causality and how we can make causal inferences (you will not be tested on designs if we do not end up covering them);
  3. your ability to think through and generate research designs to answer research questions; and
  4. your ability to interpret code, results from statistical analyses, and data visualizations.

Neither lectures nor TF sections will be explicitly focused on the exam (though obviously we will be covering material that will be on the exam in both).

ImportantThe exam will only be on class material

You will only be tested on material explicitly covered in the course. If we didn’t read about it or cover it in class then you will not be expected to know it.

Grades

ImportantThe dreaded grade capping

According to directives from the University, instructors have been advised to operate this academic year as if the new grading policy were in effect.3 That being said, the goal of this course is for as many of you as possible to learn as much as possible and to earn a, correspondingly, high grade.

3 I, and others, actually think that capping grades is not the most effective strategy to address grade inflation (in fact, it may lead to myriad unintended, negative consequences). In fact, to fully address this issue we need a major reappraisal of the purpose (and societal-level goals) of higher education (and corresponding national-level policies aimed at both enabling and pushing universities to accomplish said purpose/goals). But I digress…

Overall, we (both you as the student and I as the instructor) should both be focusing on learning as our primary outcome, not the specific grade you earn. Now while grades, ideally, should be a direct reflection of how much you have learned in a course, they are, in fact, imperfect signals. While it’s easier said than done (I was a student too not too long ago…), you should focus on learning the material as much as possible and the grades will follow.

Here’s the breakdown of how your final grade will be calculated:

Category Percent of Final Grade
Warmup Questions 10%
Class Participation 10%
Section Participation 10%
Homework 20%
Section Project 20%
Final Exam 30%
NoteBump-up policy

As mentioned above, we reserve the right to “bump up” the grades of students who significantly participate in Study Halls and on the Ed Discussion Board (this includes posting questions but, more importantly, answering questions, engaging with content, and sharing interesting things too).

Course policies

Laptops and written materials

Given that this course heavily focuses on data analysis we will be using computers consistently in class, you should always bring your laptop. That being said, we will be putting away our laptops for significant portions of class. Consequently, you will be expected to use some sort of real, existing material to take notes. This could be printer paper, a notebook or binder, even chalk-and-slate, just so long as you are able to physically write notes with some sort of writing utensil. This is not only important because you won’t have access to your laptop to take notes but also because physically writing notes is far better for memorization and retention than typing (see this study for a summary of this research).

Academic honesty

Don’t cheat! Like really, just don’t do it. We also take plagiarism (including the use of AI to generate text) very seriously and expect you to write all of your own text. If you are found to have cheated on any of your assignments (or plagiarized someone else’s work) you will be reported directly to the Office of Academic Integrity and Student Conduct.

Coding and AI

While the use of AI for coding/programming has proliferated both in industry and academia, the use of AI when learning has significant deleterious effects. You will not be able to use AI on in-section coding problem sets, though we will be discussing how to use AI in research in section.

Simply put, using AI in this course will actively hurt you in a) understanding the course material and b) preparing for assessments where you will have no access to said AI.

ImportantAI for warmup questions, homework, or any writing is expressly prohibited

As mentioned above, you are explicitly prohibited from using AI to respond to any part any warmup or homework assignments (AI is also obviously not allowed for the exam). If you are found to be using AI for the purposes of text generation you will immediately be reported to the Office of Academic Integrity and Student Conduct.

Late policy

Warmup assignments are due 1 hour prior to the start of class online and homework assignments are due by the start of class in-person. These assignments will not be accepted late under any circumstances.

Regrading policy

If you believe that there has been a grading error for your warmup questions, homework, or in-section coding problem sets and you have proof of said error then you may be granted a regrade of the error. Requests without sufficient evidence will be denied.

For the exam and final project any regrade requests will result in a complete regrade by a different TF or Dr. Fuller—the entire assignment will be graded again and will have a new grade assigned. There is no guarantee of a higher score.

Requests should be sent to the head TF Nicholas Conroy.

Office hours, study halls, and availability

Office hours and study halls for all of the teaching staff are listed on the Office Hours page.

If you have a general question, you should post it to either the class or your section Ed Discussion Board. This is by far the fastest way to get an answer, much faster than emailing one of the teaching staff. However, you can also email me directly at sfuller@fas.harvard.edu if you have a question that is about a personal situation.

Course materials

Books

As mentioned above We will use the following books in this class:

  • Regression and Other Stories by Andrew Gelman, Jennifer Hill, & Aki Vehtari

  • The Effect: An Introduction to Research Design and Causality by Nick Huntington-Klein.

Both of these books are free online but you can also purchase hard-copies if you’re a big spender (or just prefer hard copies).

Computing

We’re gonna be using R, and a whole bunch of it, to conduct our data analysis. The nice thing about R is that it’s free, open-source, and available on essentially every platform (yes, even you Linux sickos). While some of you may be wondering why we aren’t using something like Python or Stata the answer is three-fold:

  1. R has a ridiculous amount of support and capabilities when it comes to data analysis and it is overwhelmingly the most common language for Political Science (and the social sciences generally, though economists and some sociologists still love Stata for some reason; there are those in the social sciences who even use SAS or SPSS, but we don’t talk about them…)

  2. it’s completely free (Stata is obscenely expensive, though there may be a Harvard license, I haven’t checked) and

  3. it’s super easy to install and setup (yes, I’m criticizing you, Python).

We’re also going to use a ““graphical user interface”” (scare quotes intended) called RStudio to actually interact with R (if you’re familiar with Python this is like using VSCode to work with Python). All you need to know is that RStudio makes it a lot easier to code in R. So use it.

Details on how do get up and running with R and RStudio are available here

  • You can download R here, just choose your operating system.

  • You can download RStudio here, just choose your operating system.

Feel free to download these before your first section, but that meeting will cover this and more.

Accessibility

If you need reasonable accommodations for this course please contact the Disability Access Office (DAO) and ensure that I (Dr. Fuller) am notified as soon as possible. Accommodations will not alter the core requirements of the course and will not be retroactive. If something changes during the semester, you should contact DAO.

Acknowledgements

First, I want to thank Matt Blackwell for all of the materials that have served as a foundation from which this course (and site) is built. Second, I’d like to thank Jack Rametta and Chris Hare for their recommendations on books and pedagogy. Third, I especially want to thank Noah Shenker (previous head TF of the course for two years) for invaluable suggestions and feedback for how this course could be improved. And finally, I would like to thank all of the TFs and CAs (especially Head TF Nicholas Conroy) for their hard work and significant contributions to this course.