Data
My research relies on data infrastructures that make political elites observable across settings and over time. This page presents the datasets, measurement systems, and AI-assisted workflows that support the substantive papers described on the Research page. Together, these resources expand comparative coverage, scale document-based data construction, and extend political measurement beyond conventional biographical and textual records.
Comparative coverage
Comparative Political Elites Data
I am building a modular collection of political biographies that traces careers, education, political affiliations, family ties, and institutional positions across national and transnational settings. Individual modules are released separately while being aligned to a common data structure.
Chinese Political Elites
Career histories and official portraits for more than 4,000 mid- and senior-level officials, including promotion timing and purge outcomes. These data support Portraits of Power and related research on political selection.
US Congressional Legislators
A database of 8,841 senators and representatives spanning 1757–2025, with biographical information, education, career paths, party affiliations, geography, and political-family networks. Explore the database →
OECD Ministerial Officials
A comparative database of 3,642 ministerial-level officials across 36 OECD member countries, covering education, prior careers, party affiliation, ministerial portfolios, and family networks. Explore the database →
Career Trajectories of International-Organization Officials
Biographical and career data for 6,870 officials who have served in international organizations. The collection contains 69,473 career records, including 21,794 positions within international organizations as well as pre- and post-IO careers in governments, universities, NGOs, and the private sector. Explore the database →
Russia and Southeast Asia Modules
New national modules extend the collection to Russian political elites and ten Southeast Asian countries. Data cleaning and harmonization are ongoing.
These datasets are designed to support comparative research on elite recruitment and mobility; education, family background, and political networks; and movement across parties, governments, international organizations, and private institutions.
Data production
AI-Assisted Data Construction
My methodological work addresses two linked problems in document-based data production: how relevant evidence enters a dataset, and how coding rules are executed consistently once that evidence has been assembled.
Evidence Discovery and Synthesis
Agentic Framework for Political Biography Extraction develops a two-stage workflow in which agents search for and synthesize biographical evidence before an LLM coder converts the curated material into structured records. The project is under revise and resubmit as a Research Note at the American Journal of Political Science.
Codebook Execution and Governance
The Codebook Is Not a Prompt studies how complex social-science coding rules can be compiled into adaptive workflows that decompose tasks, identify relevant objects, escalate ambiguous cases to experts, preserve adjudication decisions, and apply revised rules consistently across previously coded cases. The project is in development.
Together, the two projects describe a continuous production chain: the first governs how evidence enters the system; the second governs how research rules are applied within it.
Measurement
Multimodal Measurement
Biographical records reveal where elites come from and how their careers unfold, but they do not capture how authority, conflict, and emotion are performed in political settings. I therefore use images, voice, and video to construct measures of perceived appearance, vocal delivery, and facial affect.
Official Portraits and Political Selection
The data behind Portraits of Power link official portraits to human and machine estimates of perceived competence, trustworthiness, aggressiveness, and attractiveness, as well as detailed promotion and purge outcomes for more than 4,000 Chinese officials.
Voice, Face, and Legislative Behavior
The South Korean National Assembly project combines 355 hours of plenary video, 11,060 speech segments, facial and vocal measures, transcripts, and legislator biographies. The associated paper, What Drives Legislative Anger, examines how partisan blame and changes between government and opposition shape emotional expression.
Dynamic Affect in Political Hierarchy
The Disciplined Face of Power draws on more than 20,000 video appearances of Chinese officials to study how rank and institutional role shape calmness, stress, and the management of facial affect over time.
Shared platform
Risk-A-Lab Data Intelligence
At Peking University, I lead the day-to-day development of Risk-A-Lab Data Intelligence, an interactive platform that organizes original databases, harmonized public datasets, and tools for discovery, documentation, visualization, and reuse. The platform extends this infrastructure-building work from individual research projects to shared resources for political science and international relations.
Risk-A-Lab Data Intelligence is a team-built platform. Editorial direction is set by the Lab Director; I lead day-to-day platform development, while database maintainers and research assistants contribute individual datasets and documentation.
Access
Data Access and Collaboration
Public modules can be explored through Risk-A-Lab Data Intelligence. Replication materials for published findings will be released in accordance with journal requirements. Some image, video, licensed, or personally identifying source materials cannot be redistributed; where possible, I will provide source documentation, derived measures, code, and metadata needed to understand and reproduce the data-production process.
For data access, documentation, or collaboration inquiries, please contact me at yangsp@pku.edu.cn.