Testing What We Think We Know About Civil Wars: Harmonising Explanation and Prediction in Social Science

Social science has for decades measured the quality of a model by its coefficient estimates, statistical significance and theoretical logic. If an equation explained historical data and produced low p-values, it was accepted as a proof of a causal mechanism. Yet, a model can fit historical data perfectly while being incapable of out-of-sample forecasting. Civil war research provides a clear lens into this dilemma, showing how traditional statistical approaches struggle when asked to predict the onsets of civil wars. 

Muchlinski et al. (2016) directly confronted this dilemma by examining the predictive power of political science’s most known explanatory frameworks. They looked at past literature such as The Collier-Hoeffler Model of Civil War Onset and the Case Study Project Research Design by Collier, Hoeffler, & Sambanis (2005) to evaluate if those explanatory models can predict the onset of civil wars on out-of-sample data. 

Collier, Hoeffler, & Sambanis (2005) argued in their research that civil war onset is caused by opportunity rather than grievance. In other words, low recruitment cost of young men, overseas financial diasporas and accessible terrain matter more than economic inequality, political repression, ethnic and religious diversity in the onset of civil wars. Using multivariate logistic regression, they evaluated 750 five-year country episodes from 161 countries covering years 1960 to 1999. Their in-sample model concluded that civil wars are caused by low economic opportunity cost, geographical advantages and primary commodity exports. As these variables held statistically significant coefficient estimates, social science broadly accepted them as proven causal mechanisms. 

Muchlinski et al. (2016) directly challenged the traditional explanatory approaches used by Collier, Hoeffler, & Sambanis (2005) as well as those by Fearon and Laitin (2003) and Hegre and Sambanis (2006) and questioned whether they could actually forecast civil war on unseen data. To audit these foundational studies, they conducted out-of-sample tests comparing different research models, such as standard logistic regression, Firth’s penalised logistic regression, LASSO logistic regression and Random Forest classification. The models were trained on one portion of the data and forced to predict civil war onsets on unseen data. What the authors found was that the explanatory models failed to predict rare-events, they produced massive false-negative rates and achieved the Area Under the Curve (AUC) scores between 0.70 – 0.80. While the Random Forest approach, a machine learning algorithm capable of capturing complex, non-linear interactions, achieved a much higher AUC score of 0.91. This contrast demonstrated that a high in-sample fit can suffer great predictive blindness.

While Muchlinski et al. (2016) highlighted the limits of traditional logistic regression, their results sparked a debate regarding the true size of Random Forest predictive advantage. Later papers and replications, such as Neunhoeffer & Sternberg (2018), tested the predictive advantage of Muchlinski et al’s paper and found that Muchlinski et al’s original AUC jump was partially inflated by data leakage and improper cross-validation. When models were evaluated using strict out of time validation, predicting strictly into the future rather than across randomly shuffled years, the performance difference between machine learning and traditional models narrowed significantly. Just as traditional explanatory models can hide behind misleading p-values and in-sample overfitting, predictive algorithms can hide behind inflated accuracy metrics.

This realization creates a double warning for social science: just as we cannot blindly rely on explanatory theory, we cannot simply throw massive data sets and complex algorithms at a model to replace theoretical thinking altogether. Rød, Hegre, and Leis in their paper Predicting Armed Conflict Using Protest Data (2025) provide a cautionary lesson to the purely data driven approach. The authors asked whether high frequency event data, more specifically political protest data, can improve out-of-sample forecasts of armed conflict onset. As protests represent a form of political unrest that can lead to violent escalation, it makes them perfect indicators of an early conflict warning system. Yet, they discovered that as they integrated raw, unstructured protest data directly into machine learning algorithms they performed worse at out-of-sample prediction than standard ‘baseline’ models. Without structural context, the machine learning models picked up excessive noise. Most protests do not escalate into civil wars, so purely data driven ‘naïve’ models, lacking political context, struggle to differentiate between a routine civic expression and actual pre-war mobilization, generate high false-positive rates and therefore perform worse. However, when Rød, Hegre, and Leis (2025) integrated protest indicators into theoretically informed models, including regime type, state capacity, political inclusion and conflict history, the predictive accuracy improved drastically. Their findings teach an important lesson for modern social science. Predictive models cannot operate in a theoretical vacuum. Big data and machine learning are simply amplifiers of knowledge, not substitutes for theoretical understanding. 

This is not a contest in which prediction replaces explanation. Explanatory models can look convincing until prediction exposes what they do not know; predictive models can look unbeatable until theory exposes what they do not know. Prediction and explanation should be combined as partners rather than rivals. The most rigorous social science happens at this intersection, where explanation gives prediction its meaning and prediction gives explanation its empirical validation. At CERSP, this is the intersection we are interested in.

[Written by CERSP intern Amélie Benková]

Scroll to Top