Harvard's Rating System, Decoded: What the Numbers Actually Reward
Haloway Team · · 7 min read
For decades, how elite admissions offices actually evaluated applicants was a black box, and an entire industry of speculation grew up around it. Then, through the discovery process in the SFFA v. Harvard lawsuit, the box opened: internal documents and data made public in the litigation revealed the rating system Harvard's readers used and, more valuably, the admission rates attached to each rating. It remains the clearest look inside a top admissions office that outsiders have ever gotten.
The data is historical, and I'll come back to why that caveat has teeth. But the structure it reveals is worth understanding, because it confirms some conventional wisdom, demolishes other parts of it, and puts hard numbers on things counselors usually only gesture at.
The system.
Harvard's readers rated each applicant in four dimensions (academic, extracurricular, athletic, and personal) on a scale of 1 to 6, where 1 is best and vanishingly rare. Each dimension answers a distinct question: Can this student thrive intellectually here? What have they achieved beyond the classroom? Can they contribute through athletics? And what kind of classmate and community member will they be?
The headline finding is how the ratings were distributed. A 2 was the mark of a genuinely excellent applicant: historically about 42% of applicants earned an academic 2, about 24% an extracurricular 2, about 21% a personal 2. A 1, in any category, was practically a different species: roughly 0.5% of applicants received an academic 1, 0.3% an extracurricular 1, 0.9% an athletic 1, and the personal 1 went to fewer than fifty applicants a year, effectively zero percent of the pool.
What separates a 1 from a 2.
This is the most instructive part of the whole dataset, because the gap between 1 and 2 is not "slightly better."
An academic 2 meant superb grades, demanding coursework, and test scores in the top ranges. In other words, perfect execution of the normal high school system. An academic 1 required something the normal system can't produce: original research, national or international academic distinction, published work with real intellectual contribution, evidence that the student was operating beyond the high-school level entirely. Read that again, because it's the finding that should reorganize how ambitious families think: a student with flawless grades and perfect scores still typically earned a 2. Perfection at the standard game capped out one level below the top, because tens of thousands of applicants play the standard game perfectly.
The extracurricular scale had the same structure. A 2 covered elite performance at the school or regional level: student-body president, newspaper editor, real leadership with real results. A 1 required national-level accomplishment, professional-quality work, or impact rare even inside Harvard's applicant pool. Titles alone moved nothing; the litigation documents make clear that scale, difficulty, and results were what separated the levels, not the word "founder."
The personal rating, the most controversial category in the lawsuit and the most subjective, ran on evidence from recommendations, interviews, essays, and school reports. A 2 described someone likable, mature, and clearly positive for a community. The near-mythical 1 described someone unusually memorable: warm, generous, at ease among talented peers, the person an interviewer remembers decades later. Notably, nothing in the documents equates this with charisma or extroversion; the quiet student with unmistakable integrity, empathy, and steadiness was fully eligible. What the category couldn't be was manufactured, since it was assembled almost entirely from other people's testimony about years of behavior.
The numbers that reorganize strategy.
Now the part that made the litigation data famous among admissions obsessives: the historical admission rates by rating profile.
An applicant with a single 1 and no others was admitted at roughly 68% (academic 1), 48% (extracurricular 1), 66% (personal 1), or 88% (athletic 1). Against a base rate in the low single digits, one rating of 1 in any dimension multiplied an applicant's odds by an order of magnitude or more.
Applicants with four 2s, excellent across every single dimension, were admitted at about 68%. Three 2s and a weaker fourth rating: about 43%.
And applicants with no 1s and no 2s anywhere, students who were by any normal standard strong, were admitted at approximately 0.1%. One in a thousand.
Put together, the data says there were exactly two profiles with real chances: exceptional in one dimension (a genuine 1, with everything else merely solid), or exceptional across nearly all of them (2s everywhere, which is far rarer than it sounds). Everything else, the vast, capable, hardworking middle of the pool, was competing for statistical noise.
The "well-rounded" autopsy.
This is where the data performs its most useful demolition. "Harvard wants well-rounded students" is technically supported by the four-2s profile, until you look at what a 2 actually meant. Four 2s is not a student with good grades, several clubs, a sport, an instrument, and volunteer hours. It's a student who is outstanding in four dimensions simultaneously: leadership with results, regional recognition, top-range academics, glowing personal testimony. That's not well-rounded in the guidance-office sense. That's four separate excellences occupying one teenager.
The familiar well-rounded profile, respectable everything and distinguished nothing, mostly generated 3s: solid participation, no unusual distinction. And the 0.1% figure is the data's verdict on that profile. The question the ratings asked was never "how many things did this student do?" It was "at what level, and what changed because they did it?"
The caveats, which are real.
I want to be as clear about the limits of this data as about its lessons, because it's routinely abused as a calculator.
These figures come from litigation over admissions cycles that are now years in the past. Since then: the Supreme Court's 2023 ruling ended race-conscious admissions and pushed offices to revise practices; testing policies have convulsed twice; the applicant pool has grown and shifted; and Harvard has every institutional reason to have altered a rating system whose details were dragged into federal court. The percentages describe groups, historically; they were never individual probabilities, and a 68% admit rate still meant a third of those extraordinary applicants were rejected, for reasons of institutional priorities, class composition, and space that no rating captures. Context also ran through everything: readers evaluated achievement against available opportunity, and the rating scheme itself had categories acknowledging students whose "activities" were employment and caregiving.
So treat this as an X-ray of how one elite office structured judgment, not as this year's formula at any school.
What survives the caveats.
Strip away the specific percentages and the durable lessons are these, and they're worth internalizing precisely because they came from internal documents rather than admissions-office PR:
Conventional perfection is a 2, not a 1. The standard system, executed flawlessly, produces an excellent-but-common profile. The top rating in every category required going outside the system (original work, external recognition, real scale), which is exactly the kind of achievement that takes years and can't be assembled at the end.
Depth had a measurable exchange rate. One genuine distinction moved admission odds more than any accumulation of memberships. The data doesn't merely suggest that depth beats breadth-as-checklist; it prices the difference.
The personal dimension carried real weight and resisted engineering. A personal 1 rivaled an academic 1 in its effect on admission, and it was built from years of other people's observations. Reverse-engineering a personality for applications fails on its own terms; the only reliable strategy for that category is the suspiciously simple one of actually being, over a long period, someone teachers and peers genuinely want to vouch for.
And nothing guaranteed anything. The strongest profiles in the dataset still faced meaningful rejection rates. The ratings describe how excellence was recognized, not how certainty was purchased, because certainty was never for sale.
Which points, as this data always does when you sit with it long enough, away from the ratings themselves. You can't apply to a rating. You can only spend high school becoming the kind of person the ratings were built to detect: genuinely excellent at something real, strong enough everywhere else, and, on the years-long testimony of the people who actually know you, good to have in the room.
An admissions copilot and workspace. Built to help make your dreams come true.
Worker smarter, apply stronger. See how Haloway boost your application and catches mistakes before you submit.