Lezing
SAILS Lunch Time Seminar: Matthew di Giuseppe
- Datum
- maandag 26 oktober 2026
- Tijd
- Locatie
- Online only
Assessing Superhuman LLM Capabilities for Social Science Measurement by Cross-Sectional Prediction
Abstract:
Social scientists increasingly ask LLMs to perform measurement tasks no human coder could complete. For example, placing 702 occupations on a scale of exposure to digitization or scoring the relative political ideology of hundreds of international political parties. The discipline's validation standard, agreement with trained annotators, cannot be applied there, because the annotators who would supply the benchmark do not exist. Further, any benchmark whose answer was already knowable may sit inside the training corpus, so a passing capability check is indistinguishable from a memory test.
I propose a partial solution built on cross-sectional prediction. The design rests in the gap between model training and public release and that shocks known to researchers are unknown to models. I treat the model's domain knowledge as a construct in its own right, select an domain relevant event whose outcomes were realized after the model's training cutoff, and have the model place a frozen pre-event cross-section of entities on a latent dimension of exposure to it. The realized outcomes then supply the ordering against which the measure is scored. The existing out-of-training benchmarks score independent questions one at a time and aggregate by a proper scoring rule. An ordering of hundreds of heterogeneous units under a single shock asks a harder question, whether the model can weigh competing transmission channels against each other across units that differ on all of them at once. Because a common shock grades every unit rather than a sampled handful, the misses are useful. A follow-up analysis separates residuals driven by changing context from those that are legitimate misses.
I demonstrate this by scaling 503 S&P 500 firms on vulnerability to a closure of the Strait of Hormuz, aggregating 7,446 pairwise model comparisons with a Bradley-Terry model, and grading the ranking against realized excess returns after the strait closed in February 2026, six months past the model's published knowledge cutoff. The ordering tracks the realized cross-section at short horizons, survives fixed effects for 99 sub-industry categories, and beats five mechanical benchmarks built from pre-event information alone. Further, it predicts just as well for the obvious cases as it does for the middle 60% of the scale. The approach is imperfect but likely the only way to assess LLM capabilities for super human tasks.
Join us!
The SAILS Lunch Time Seminar is an online event, but it is not publicly accessible in real-time. Please click the the link below to register to our mailinglist and receive participation links for our Lunch Time Seminars.
Sign up