References
Abdulkadiroğlu, Atila, Joshua D Angrist, Yusuke Narita, and Parag A
Pathak. 2017. “Research Design Meets Market Design: Using
Centralized Assignment for Impact Evaluation.”
Econometrica 85 (5): 1373–432.
Antonakis, John, Nicolas Bastardoz, and Mikko Rönkkö. 2021. “On
Ignoring the Random Effects Assumption in Multilevel Models: Review,
Critique, and Recommendations.” Organizational Research
Methods 24 (2): 443–83.
Benjamin, Daniel J., James O. Berger, Magnus Johannesson, et al. 2017.
“Redefine Statistical Significance.” Nature Human
Behaviour 2 (1): 6–10. https://doi.org/10.1038/s41562-017-0189-z.
Bloom, Howard S., Stephen W. Raudenbush, Michael J. Weiss, and Kristin
Porter. 2016. “Using Multisite Experiments to
Study Cross-Site Variation in Treatment Effects: A Hybrid Approach With
Fixed Intercepts and a Random Treatment Coefficient.”
Journal of Research on Educational Effectiveness 10 (4): 0–0.
https://doi.org/10.1080/19345747.2016.1264518.
Boos, Dennis D., and Jason A. Osborne. 2015. “Assessing
Variability of Complex Descriptive Statistics in Monte
Carlo Studies Using Resampling Methods.” International
Statistical Review 83 (2): 228–38. https://doi.org/10.1111/insr.12087.
Borenstein, Michael, Larry V. Hedges, Julian P. T. Higgins, and Hannah
R. Rothstein. 2021. Introduction to Meta-Analysis. 3rd ed. John
Wiley & Sons.
Boulesteix, Anne-Laure, Sabine Hoffmann, Alethea Charlton, and Heidi
Seibold. 2020. “A Replication Crisis in Methodological
Research?” Significance 17 (5): 18–21. https://doi.org/10.1111/1740-9713.01444.
Boulesteix, Anne-Laure, Sabine Lauer, and Manuel J. A. Eugster. 2013.
“A Plea for Neutral Comparison Studies in Computational
Sciences.” PLOS ONE 8 (4): e61562. https://doi.org/10.1371/journal.pone.0061562.
Boulesteix, Anne-Laure, Rory Wilson, and Alexander Hapfelmeier. 2017.
“Towards Evidence-Based Computational Statistics: Lessons from
Clinical Research on the Role and Design of Real-Data Benchmark
Studies.” BMC Medical Research Methodology 17 (1, 1):
1–12. https://doi.org/10.1186/s12874-017-0417-2.
Brown, Morton B., and Alan B. Forsythe. 1974. “The Small
Sample Behavior of Some Statistics Which Test the
Equality of Several Means.”
Technometrics 16 (1): 129–32. https://doi.org/10.1080/00401706.1974.10489158.
Brown, M, H Solomon, and M A Stephens. 1977. “Estimation of
Parameters of Zero-One Processes by Interval Sampling.”
Operations Research 25 (3): 493–505.
Cameron, A Colin, and Douglas L Miller. 2015. “A Practitioner’s
Guide to Cluster-Robust Inference.” Journal of Human
Resources 50 (2): 317–72. https://doi.org/10.3368/jhr.50.2.317.
Chen, Man, and James E. Pustejovsky. 2024. “Multi-Level
Meta-Analysis of Single-Case Experimental Designs Using Robust Variance
Estimation.” Psychological Methods 29 (3): 537–60. https://doi.org/10.1037/met0000510.
Chen, Man, and James E. Pustejovsky. 2025. “Adapting Methods for
Correcting Selective Reporting Bias in Meta-Analysis of Dependent Effect
Sizes.” Psychological Methods, ahead of print, June 9.
https://doi.org/10.1037/met0000773.
Cho, Hunyong, Chuwen Liu, John S Preisser, and Di Wu. 2023. “A
Bivariate Zero-Inflated Negative Binomial Model and Its Applications to
Biomedical Settings.” Statistical Methods in Medical
Research 32 (7): 1300–1317. https://doi.org/10.1177/09622802231172028.
Davison, A. C., and D. V. Hinkley. 1997. Bootstrap Methods and Their
Applications. Cambridge University Press.
Dong, Nianbo, and Rebecca Maynard. 2013.
“PowerUp! : A
Tool for Calculating Minimum Detectable Effect Sizes
and Minimum Required Sample Sizes for
Experimental and Quasi-Experimental Design
Studies.” Journal of Research on Educational
Effectiveness 6 (1): 24–67. https://doi.org/10.1080/19345747.2012.673143.
Dorie, Vincent, Jennifer Hill, Uri Shalit, Marc Scott, and Dan Cervone.
2019. “Automated Versus Do-It-Yourself Methods for Causal
Inference: Lessons Learned from a Data Analysis Competition.”
Statistical Science 34 (1): 43–68. https://doi.org/10.1214/18-STS667.
Efron, Bradley. 2000. “The Bootstrap and Modern
Statistics.” Journal of the American Statistical
Association 95 (452): 1293–96. https://doi.org/10.2307/2669773.
Efron, Bradley, and Robert J Tibshirani. 1994. An Introduction to
the Bootstrap. Chapman; Hall/CRC.
Enders, Craig K, Brian T Keller, and Michael P Woller. 2023. “A
Simple Monte Carlo Method for Estimating Power in Multilevel
Designs.” Psychological Methods.
Faul, Franz, Edgar Erdfelder, Axel Buchner, and Albert-Georg Lang. 2009.
“Statistical Power Analyses Using
G*Power 3.1: Tests for
Correlation and Regression Analyses.” Behavior Research
Methods 41 (4): 1149–60. https://doi.org/10.3758/BRM.41.4.1149.
Franklin, Jessica M., Sebastian Schneeweiss, Jennifer M. Polinski, and
Jeremy A. Rassen. 2014. “Plasmode Simulation for the Evaluation of
Pharmacoepidemiologic Methods in Complex Healthcare Databases.”
Computational Statistics & Data Analysis 72: 219–26. https://doi.org/10.1016/j.csda.2013.10.018.
Freedman, David A. 2008. “On Regression Adjustments to
Experimental Data.” Advances in Applied Mathematics 40
(2): 180–93.
Friedman, Linda Weiser, and Israel Pressman. 1988. “The Metamodel
in Simulation Analysis: Can It Be Trusted?” Journal of the
Operational Research Society 39 (10): 939–48.
Fryda, Tomas, Erin LeDell, Navdeep Gill, et al. 2014. H2o: R
Interface for the ’H2O’ Scalable Machine
Learning Platform. Comprehensive R Archive Network. https://doi.org/10.32614/CRAN.package.h2o.
Gelman, Andrew, John B. Carlin, Hal S. Stern, David B. Dunson, Aki
Vehtari, and Donald B. Rubin. 2013. Bayesian Data
Analysis. Chapman and Hall/CRC. https://doi.org/10.1201/b16018.
Gelman, Andrew, Cristian Pasarica, and Rahul Dodhia. 2002. “Let’s
Practice What We Preach: Turning Tables into Graphs.” The
American Statistician 56 (2): 121–30.
Gentzel, Amanda M., Purva Pruthi, and David Jensen. 2021. “How and
Why to Use Experimental Data to Evaluate Methods for Observational
Causal Inference.” Proceedings of the 38th International
Conference on Machine Learning, Proceedings of machine learning
research, vol. 139: 3660–71. https://proceedings.mlr.press/v139/gentzel21a.html.
Gerber, Alan S, and Donald P Green. 2017. “Field Experiments on
Voter Mobilization: An Overview of a Burgeoning Literature.”
Handbook of Economic Field Experiments 1: 395–438.
Gilbert, Joshua, and Luke Miratrix. 2024. “Multilevel Metamodels:
A Novel Approach to Enhance Efficiency and Generalizability in Monte
Carlo Simulation Studies.” arXiv Preprint
arXiv:2401.07294.
Good, Phillip. 2013. Permutation Tests: A Practical Guide to
Resampling Methods for Testing Hypotheses. Springer Science &
Business Media.
Hettmansperger, Thomas P., and Joseph W. McKean. 2010. Robust
Nonparametric Statistical Methods. 2nd ed. Chapman &
Hall / CRC Monographs on
Statistics & Applied Probability.
Taylor and Francis.
Hill, Jennifer L. 2011. “Bayesian Nonparametric Modeling for
Causal Inference.” Journal of Computational and Graphical
Statistics 20 (1): 217–40. https://doi.org/10.1198/jcgs.2010.08162.
Hoogland, Jeffrey J., and Anne Boomsa. 1998. “Robustness
Studies in Covariance Structure Modeling:
An Overview and a Meta-Analysis.”
Sociological Methods & Research 26 (3): 329–67. https://doi.org/10.1177/0049124198026003003.
Hunter, Kristen B., Luke Miratrix, and Kristin Porter. 2024.
“PUMP: Estimating Power, Minimum Detectable Effect Size, and
Sample Size When Adjusting for Multiple Outcomes in Multi-Level
Experiments.” Journal of Statistical Software 108 (6):
1–43. https://doi.org/10.18637/jss.v108.i06.
James, G. S. 1951. “The Comparison of Several Groups of
Observations When the Ratios of the Population Variances Are
Unknown.” Biometrika 38 (3/4): 324. https://doi.org/10.2307/2332578.
Jones, Owen, Robert Maillardet, and Andrew Robinson. 2012.
Introduction to Scientific Programming and
Simulation Using R. Chapman and Hall/CRC.
https://doi.org/10.1201/9781420068740.
Kang, Hyunseung, Laura Peck, and Luke Keele. 2018. “Inference for
Instrumental Variables: A Randomization Inference Approach.”
Journal of the Royal Statistical Society Series A: Statistics in
Society 181 (4): 1231–54.
Kern, Holger L., Elizabeth A. Stuart, Jennifer Hill, and Donald P.
Green. 2014. “Assessing Methods for
Generalizing Experimental Impact Estimates to Target
Populations.” Journal of Research on Educational
Effectiveness 9 (1): 103–27. https://doi.org/10.1080/19345747.2015.1060282.
Kleijnen, Jack PC. 1981. “Regression Analysis for Simulation
Practitioners.” Journal of the Operational Research
Society 32 (1): 35–43.
Lakens, Daniel, Federico G. Adolfi, Casper J. Albers, et al. 2018.
“Justify Your Alpha.” Nature Human Behaviour 2
(3): 168–71. https://doi.org/10.1038/s41562-018-0311-x.
Lee, Young Ri, and James E Pustejovsky. 2023. “Comparing Random
Effects Models, Ordinary Least Squares, or Fixed Effects with Cluster
Robust Standard Errors for Cross-Classified Data.”
Psychological Methods.
Lehmann, Erich Leo et al. 1975.
“Statistical Methods Based on Ranks.” Nonparametrics.
San Francisco, CA, Holden-Day 2.
Lin, Winston. 2013. “Agnostic Notes on Regression Adjustments to
Experimental Data: Reexamining Freedman’s Critique.” The
Annals of Applied Statistics, 295–318.
Little, Roderick J. 2013. “In Praise of Simplicity Not
Mathematistry! Ten Simple Powerful Ideas for the Statistical
Scientist.” Journal of the American Statistical
Association 108 (502): 359–69.
Long, J. Scott, and Laurie H. Ervin. 2000. “Using
Heteroscedasticity Consistent Standard Errors in the Linear Regression
Model.” The American Statistician 54 (3): 217–24. https://doi.org/10.1080/00031305.2000.10474549.
Maechler, Martin, Peter Rousseeuw, Christophe Croux, et al. 2024.
Robustbase: Basic Robust Statistics. http://robustbase.r-forge.r-project.org/.
Maronna, Ricardo A., R. Douglas Martin, and Vı́ctor J. Yohai. 2006.
Robust Statistics: Theory and Methods. Wiley Series in
Probability and Statistics. J. Wiley.
McKean, Joseph W., and Ronald M. Schrader. 1984. “A Comparison of
Methods for Studentizing the Sample Median.” Communications
in Statistics - Simulation and Computation 13 (6): 751–73. https://doi.org/10.1080/03610918408812413.
Mehrotra, Devan V. 1997. “Improving the Brown-Forsythe Solution to
the Generalized Behrens-Fisher Problem.” Communications in
Statistics - Simulation and Computation 26 (3): 1139–45. https://doi.org/10.1080/03610919708813431.
Miratrix, Luke W., Michael J. Weiss, and Brit Henderson. 2021. “An
Applied Researcher’s Guide to Estimating
Effects from Multisite Individually Randomized
Trials: Estimands, Estimators, and
Estimates.” Journal of Research on Educational
Effectiveness 14 (1): 270–308. https://doi.org/10.1080/19345747.2020.1831115.
Morris, Tim P., Ian R. White, and Michael J. Crowther. 2019.
“Using Simulation Studies to Evaluate Statistical Methods.”
Statistics in Medicine, ahead of print, January. https://doi.org/10.1002/sim.8086.
Oehlert, Gary W. 1992. “A Note on the Delta
Method.” The American Statistician 46 (1): 27–29.
https://doi.org/10.1080/00031305.1992.10475842.
Pashley, Nicole E, Luke Keele, and Luke W Miratrix. 2024.
“Improving Instrumental Variable Estimators with
Poststratification.” Journal of the Royal Statistical Society
Series A: Statistics in Society, qnae073.
Pedersen, Thomas Lin. 2024. Patchwork: The Composer of Plots.
https://patchwork.data-imaginist.com.
Pustejovsky, James E. 2014. “Converting from d to r to z When the
Design Uses Extreme Groups, Dichotomization, or Experimental
Control.” Psychological Methods 19 (1): 92.
Pustejovsky, James E. 2018. The Multivariate Delta Method. https://jepusto.com/posts/Multivariate-delta-method/.
Robert, Christian, and George Casella. 2010. Introducing Monte
Carlo Methods with R. Springer. https://doi.org/10.1007/978-1-4419-1576-4.
Rosenbaum, Paul R. 2017. Observation and Experiment: An Introduction
to Causal Inference. Harvard University Press.
Rousseeuw, Peter J., and Christophe Croux. 1993. “Alternatives to
the Median Absolute Deviation.” Journal of the
American Statistical Association 88 (424): 1273–83. https://doi.org/10.1080/01621459.1993.10476408.
Satterthwaite, F. E. 1946. “An Approximate Distribution of
Estimates of Variance Components.” Biometrics Bulletin 2
(6): 110. https://doi.org/10.2307/3002019.
Shaw, Pamela A., Susan Gruber, Brian D. Williamson, et al. 2025. A
Cautionary Note for Plasmode Simulation Studies in the Setting of Causal
Inference. https://arxiv.org/abs/2504.11740.
Siepe, Björn S., František Bartoš, Tim P. Morris, Anne-Laure Boulesteix,
Daniel W. Heck, and Samuel Pawel. 2024. “Simulation Studies for
Methodological Research in Psychology: A Standardized Template for
Planning, Preregistration, and Reporting.” Psychological
Methods, ahead of print. https://doi.org/10.1037/met0000695.
Staiger, Douglas O, and Jonah E Rockoff. 2010. “Searching for
Effective Teachers with Imperfect Information.” Journal of
Economic Perspectives 24 (3): 97–118.
Tipton, Elizabeth. 2013. “Stratified Sampling Using Cluster
Analysis: A Sample Selection Strategy for Improved Generalizations from
Experiments.” Evaluation Review 37 (2): 109–39. https://doi.org/10.1177/0193841X13516324.
Tipton, Elizabeth, and James E Pustejovsky. 2015. “Small-Sample
Adjustments for Tests of Moderators and Model Fit Using Robust Variance
Estimation in Meta-Regression.” Journal of Educational and
Behavioral Statistics 40 (6): 604–34.
Tufte, Edward R, and Peter R Graves-Morris. 1983. The Visual Display
of Quantitative Information. Graphics press Cheshire, CT.
Vevea, Jack L, and Larry V Hedges. 1995. “A General Linear Model
for Estimating Effect Size in the Presence of Publication Bias.”
Psychometrika 60 (3): 419–35. https://doi.org/10.1007/BF02294384.
Welch, B. L. 1951. “On the Comparison of Several Mean Values:
An Alternative Approach.” Biometrika 38
(3/4): 330. https://doi.org/10.2307/2332579.
Westfall, Peter H, and Kevin SS Henning. 2013. Understanding
Advanced Statistical Methods. Vol. 543. CRC Press Boca Raton, FL.
White, Halbert. 1980. “A Heteroskedasticity-Consistent Covariance
Matrix Estimator and a Direct Test for Heteroskedasticity.”
Econometrica 48 (4): 817–38.
Wickham, Hadley. 2014. “Tidy Data.” Journal of
Statistical Software 59 (10): 1–23. https://doi.org/10.18637/jss.v059.i10.
Wickham, Hadley, and Jennifer Bryan. 2023. R
Packages. "O’Reilly Media, Inc." https://r-pkgs.org/.
Wickham, Hadley, Mine Çetinkaya-Rundel, and Garrett Grolemund. 2023.
R for Data Science: Import,
Tidy, Transform, Visualize, and
Model Data. "O’Reilly Media, Inc." https://r4ds.hadley.nz/.
Wilcox, Rand R. 2022. Introduction to Robust Estimation and
Hypothesis Testing. Fifth edition. Academic Press, an imprint of
Elsevier.
Zhou, Jie, Zhiwei Zhang, Zhaohai Li, and Jun Zhang. 2015.
“Coarsened Propensity Scores and Hybrid Estimators for Missing
Data and Causal Inference.” International Statistical
Review 83 (3): 449–71.