I’d like to share a small Python library and ask for feedback from people who run agent-based models more than I do.
The problem. I built Omran, an agent-based model of Ibn Khaldun’s theory of asabiyyah: nations on a grid that grow, starve, fight over borders and collapse. It produced numbers, and I couldn’t tell which of them described the model and which described the seed.
What Rasad does. You describe one run as a function, and Rasad runs it across seeds:
import rasad
def run(params, seed):
... # one run of your model
return {"final_population": ..., "population_trace": [...]}
report = rasad.measure(run, params={}, runs=100)
print(report.summary())
report.plot().write_html("divergence.html")
For each scalar output it reports the mean with its standard error and a 90% Student-t interval, then where a single run lands: std, coefficient of variation, p05 to p95 range, and a low/moderate/high label. The label’s cutoffs are a stated convention, printed in every report and set by the caller. For time series it draws a divergence curve, the standard deviation across runs at each step.
What it found in Omran (100 runs of 100 simulated years, model unmodified):
- Three of four outputs had high variability. Final population ranged from 2,704 to 6,774 across 90% of runs.
- Famines were zero in every run. The mechanism never fires.
- The same seed gave different results in different processes. Runs were identical through year 19 and split at year 20 by one individual. Running the model’s own entry point eight times with its fixed seed gave four different results. My unconfirmed suspect is iteration over a frozenset of nations, whose order follows memory addresses.
On a SimPy machine shop (10 machines, random breakdowns, one repairman), total output was stable (cv 0.014) while the gap between the best and worst machine was noise (cv 0.31).
Against the literature. Lee et al. (2015, JASSS 18(4) 4) found run counts in ABM studies that were too low, conveniently chosen or exorbitantly high, with 100 or fewer common. Rasad defaults to 100 as well, and it does not yet choose the run count by variance stability as that paper recommends. What it does is report the standard error, so you can judge whether the runs you made support the claim you want to make.
Known limitations
- Only the seed varies. Parameter sensitivity, the question SALib answers, is not implemented.
- Runs are sequential.
- The 90% interval is Student-t. I checked its coverage by simulation: 89 to 90% on normal data and 88 to 90% on Poisson data at every run count I tried. On skewed outputs it undercovers with few runs: about 83 to 85% for exponential data and 76 to 80% for lognormal (σ=1) at 5 to 10 runs, rising to 87 to 89% at 100.
- Python models only. It has been tried on Mesa, SimPy, EoN and plain Python models.
The specification and tests were written first, and AI models wrote most of the implementation against them, under review. The README documents who did what.
Questions for this community
- How do you decide how many runs to make today?
- Would a stopping rule based on variance stability, as in Lee et al., be more useful to you than a fixed run count?
- Is anything in the report misleading or missing for the way you analyse output?
Code: GitHub - ahmedaly0904-bit64/Rasad-Sim: Measure how much of a stochastic simulation's result is a property of the model and how much is the random seed. · GitHub
Install: pip install rasad-sim
