References

Abdulkadiroğlu, Atila, Joshua D Angrist, Yusuke Narita, and Parag A Pathak. 2017. “Research Design Meets Market Design: Using Centralized Assignment for Impact Evaluation.” Econometrica 85 (5): 1373–432.
Antonakis, John, Nicolas Bastardoz, and Mikko Rönkkö. 2021. “On Ignoring the Random Effects Assumption in Multilevel Models: Review, Critique, and Recommendations.” Organizational Research Methods 24 (2): 443–83.
Benjamin, Daniel J., James O. Berger, Magnus Johannesson, et al. 2017. “Redefine Statistical Significance.” Nature Human Behaviour 2 (1): 6–10. https://doi.org/10.1038/s41562-017-0189-z.
Bloom, Howard S., Stephen W. Raudenbush, Michael J. Weiss, and Kristin Porter. 2016. Using Multisite Experiments to Study Cross-Site Variation in Treatment Effects: A Hybrid Approach With Fixed Intercepts and a Random Treatment Coefficient.” Journal of Research on Educational Effectiveness 10 (4): 0–0. https://doi.org/10.1080/19345747.2016.1264518.
Boos, Dennis D., and Jason A. Osborne. 2015. “Assessing Variability of Complex Descriptive Statistics in Monte Carlo Studies Using Resampling Methods.” International Statistical Review 83 (2): 228–38. https://doi.org/10.1111/insr.12087.
Borenstein, Michael, Larry V. Hedges, Julian P. T. Higgins, and Hannah R. Rothstein. 2021. Introduction to Meta-Analysis. 3rd ed. John Wiley & Sons.
Boulesteix, Anne-Laure, Sabine Hoffmann, Alethea Charlton, and Heidi Seibold. 2020. “A Replication Crisis in Methodological Research?” Significance 17 (5): 18–21. https://doi.org/10.1111/1740-9713.01444.
Boulesteix, Anne-Laure, Sabine Lauer, and Manuel J. A. Eugster. 2013. “A Plea for Neutral Comparison Studies in Computational Sciences.” PLOS ONE 8 (4): e61562. https://doi.org/10.1371/journal.pone.0061562.
Boulesteix, Anne-Laure, Rory Wilson, and Alexander Hapfelmeier. 2017. “Towards Evidence-Based Computational Statistics: Lessons from Clinical Research on the Role and Design of Real-Data Benchmark Studies.” BMC Medical Research Methodology 17 (1, 1): 1–12. https://doi.org/10.1186/s12874-017-0417-2.
Brown, Morton B., and Alan B. Forsythe. 1974. “The Small Sample Behavior of Some Statistics Which Test the Equality of Several Means.” Technometrics 16 (1): 129–32. https://doi.org/10.1080/00401706.1974.10489158.
Brown, M, H Solomon, and M A Stephens. 1977. “Estimation of Parameters of Zero-One Processes by Interval Sampling.” Operations Research 25 (3): 493–505.
Cameron, A Colin, and Douglas L Miller. 2015. “A Practitioner’s Guide to Cluster-Robust Inference.” Journal of Human Resources 50 (2): 317–72. https://doi.org/10.3368/jhr.50.2.317.
Chen, Man, and James E. Pustejovsky. 2024. “Multi-Level Meta-Analysis of Single-Case Experimental Designs Using Robust Variance Estimation.” Psychological Methods 29 (3): 537–60. https://doi.org/10.1037/met0000510.
Chen, Man, and James E. Pustejovsky. 2025. “Adapting Methods for Correcting Selective Reporting Bias in Meta-Analysis of Dependent Effect Sizes.” Psychological Methods, ahead of print, June 9. https://doi.org/10.1037/met0000773.
Cho, Hunyong, Chuwen Liu, John S Preisser, and Di Wu. 2023. “A Bivariate Zero-Inflated Negative Binomial Model and Its Applications to Biomedical Settings.” Statistical Methods in Medical Research 32 (7): 1300–1317. https://doi.org/10.1177/09622802231172028.
Davison, A. C., and D. V. Hinkley. 1997. Bootstrap Methods and Their Applications. Cambridge University Press.
Dong, Nianbo, and Rebecca Maynard. 2013. PowerUp! : A Tool for Calculating Minimum Detectable Effect Sizes and Minimum Required Sample Sizes for Experimental and Quasi-Experimental Design Studies.” Journal of Research on Educational Effectiveness 6 (1): 24–67. https://doi.org/10.1080/19345747.2012.673143.
Dorie, Vincent, Jennifer Hill, Uri Shalit, Marc Scott, and Dan Cervone. 2019. “Automated Versus Do-It-Yourself Methods for Causal Inference: Lessons Learned from a Data Analysis Competition.” Statistical Science 34 (1): 43–68. https://doi.org/10.1214/18-STS667.
Efron, Bradley. 2000. “The Bootstrap and Modern Statistics.” Journal of the American Statistical Association 95 (452): 1293–96. https://doi.org/10.2307/2669773.
Efron, Bradley, and Robert J Tibshirani. 1994. An Introduction to the Bootstrap. Chapman; Hall/CRC.
Enders, Craig K, Brian T Keller, and Michael P Woller. 2023. “A Simple Monte Carlo Method for Estimating Power in Multilevel Designs.” Psychological Methods.
Faul, Franz, Edgar Erdfelder, Axel Buchner, and Albert-Georg Lang. 2009. “Statistical Power Analyses Using G*Power 3.1: Tests for Correlation and Regression Analyses.” Behavior Research Methods 41 (4): 1149–60. https://doi.org/10.3758/BRM.41.4.1149.
Franklin, Jessica M., Sebastian Schneeweiss, Jennifer M. Polinski, and Jeremy A. Rassen. 2014. “Plasmode Simulation for the Evaluation of Pharmacoepidemiologic Methods in Complex Healthcare Databases.” Computational Statistics & Data Analysis 72: 219–26. https://doi.org/10.1016/j.csda.2013.10.018.
Freedman, David A. 2008. “On Regression Adjustments to Experimental Data.” Advances in Applied Mathematics 40 (2): 180–93.
Friedman, Linda Weiser, and Israel Pressman. 1988. “The Metamodel in Simulation Analysis: Can It Be Trusted?” Journal of the Operational Research Society 39 (10): 939–48.
Fryda, Tomas, Erin LeDell, Navdeep Gill, et al. 2014. H2o: R Interface for the ’H2OScalable Machine Learning Platform. Comprehensive R Archive Network. https://doi.org/10.32614/CRAN.package.h2o.
Gelman, Andrew, John B. Carlin, Hal S. Stern, David B. Dunson, Aki Vehtari, and Donald B. Rubin. 2013. Bayesian Data Analysis. Chapman and Hall/CRC. https://doi.org/10.1201/b16018.
Gelman, Andrew, Cristian Pasarica, and Rahul Dodhia. 2002. “Let’s Practice What We Preach: Turning Tables into Graphs.” The American Statistician 56 (2): 121–30.
Gentzel, Amanda M., Purva Pruthi, and David Jensen. 2021. “How and Why to Use Experimental Data to Evaluate Methods for Observational Causal Inference.” Proceedings of the 38th International Conference on Machine Learning, Proceedings of machine learning research, vol. 139: 3660–71. https://proceedings.mlr.press/v139/gentzel21a.html.
Gerber, Alan S, and Donald P Green. 2017. “Field Experiments on Voter Mobilization: An Overview of a Burgeoning Literature.” Handbook of Economic Field Experiments 1: 395–438.
Gilbert, Joshua, and Luke Miratrix. 2024. “Multilevel Metamodels: A Novel Approach to Enhance Efficiency and Generalizability in Monte Carlo Simulation Studies.” arXiv Preprint arXiv:2401.07294.
Good, Phillip. 2013. Permutation Tests: A Practical Guide to Resampling Methods for Testing Hypotheses. Springer Science & Business Media.
Hettmansperger, Thomas P., and Joseph W. McKean. 2010. Robust Nonparametric Statistical Methods. 2nd ed. Chapman & Hall / CRC Monographs on Statistics & Applied Probability. Taylor and Francis.
Hill, Jennifer L. 2011. “Bayesian Nonparametric Modeling for Causal Inference.” Journal of Computational and Graphical Statistics 20 (1): 217–40. https://doi.org/10.1198/jcgs.2010.08162.
Hoogland, Jeffrey J., and Anne Boomsa. 1998. “Robustness Studies in Covariance Structure Modeling: An Overview and a Meta-Analysis.” Sociological Methods & Research 26 (3): 329–67. https://doi.org/10.1177/0049124198026003003.
Hunter, Kristen B., Luke Miratrix, and Kristin Porter. 2024. “PUMP: Estimating Power, Minimum Detectable Effect Size, and Sample Size When Adjusting for Multiple Outcomes in Multi-Level Experiments.” Journal of Statistical Software 108 (6): 1–43. https://doi.org/10.18637/jss.v108.i06.
James, G. S. 1951. “The Comparison of Several Groups of Observations When the Ratios of the Population Variances Are Unknown.” Biometrika 38 (3/4): 324. https://doi.org/10.2307/2332578.
Jones, Owen, Robert Maillardet, and Andrew Robinson. 2012. Introduction to Scientific Programming and Simulation Using R. Chapman and Hall/CRC. https://doi.org/10.1201/9781420068740.
Kang, Hyunseung, Laura Peck, and Luke Keele. 2018. “Inference for Instrumental Variables: A Randomization Inference Approach.” Journal of the Royal Statistical Society Series A: Statistics in Society 181 (4): 1231–54.
Kern, Holger L., Elizabeth A. Stuart, Jennifer Hill, and Donald P. Green. 2014. Assessing Methods for Generalizing Experimental Impact Estimates to Target Populations.” Journal of Research on Educational Effectiveness 9 (1): 103–27. https://doi.org/10.1080/19345747.2015.1060282.
Kleijnen, Jack PC. 1981. “Regression Analysis for Simulation Practitioners.” Journal of the Operational Research Society 32 (1): 35–43.
Lakens, Daniel, Federico G. Adolfi, Casper J. Albers, et al. 2018. “Justify Your Alpha.” Nature Human Behaviour 2 (3): 168–71. https://doi.org/10.1038/s41562-018-0311-x.
Lee, Young Ri, and James E Pustejovsky. 2023. “Comparing Random Effects Models, Ordinary Least Squares, or Fixed Effects with Cluster Robust Standard Errors for Cross-Classified Data.” Psychological Methods.
Lehmann, Erich Leo et al. 1975. “Statistical Methods Based on Ranks.” Nonparametrics. San Francisco, CA, Holden-Day 2.
Lin, Winston. 2013. “Agnostic Notes on Regression Adjustments to Experimental Data: Reexamining Freedman’s Critique.” The Annals of Applied Statistics, 295–318.
Little, Roderick J. 2013. “In Praise of Simplicity Not Mathematistry! Ten Simple Powerful Ideas for the Statistical Scientist.” Journal of the American Statistical Association 108 (502): 359–69.
Long, J. Scott, and Laurie H. Ervin. 2000. “Using Heteroscedasticity Consistent Standard Errors in the Linear Regression Model.” The American Statistician 54 (3): 217–24. https://doi.org/10.1080/00031305.2000.10474549.
Maechler, Martin, Peter Rousseeuw, Christophe Croux, et al. 2024. Robustbase: Basic Robust Statistics. http://robustbase.r-forge.r-project.org/.
Maronna, Ricardo A., R. Douglas Martin, and Vı́ctor J. Yohai. 2006. Robust Statistics: Theory and Methods. Wiley Series in Probability and Statistics. J. Wiley.
McKean, Joseph W., and Ronald M. Schrader. 1984. “A Comparison of Methods for Studentizing the Sample Median.” Communications in Statistics - Simulation and Computation 13 (6): 751–73. https://doi.org/10.1080/03610918408812413.
Mehrotra, Devan V. 1997. “Improving the Brown-Forsythe Solution to the Generalized Behrens-Fisher Problem.” Communications in Statistics - Simulation and Computation 26 (3): 1139–45. https://doi.org/10.1080/03610919708813431.
Miratrix, Luke W., Michael J. Weiss, and Brit Henderson. 2021. “An Applied Researcher’s Guide to Estimating Effects from Multisite Individually Randomized Trials: Estimands, Estimators, and Estimates.” Journal of Research on Educational Effectiveness 14 (1): 270–308. https://doi.org/10.1080/19345747.2020.1831115.
Morris, Tim P., Ian R. White, and Michael J. Crowther. 2019. “Using Simulation Studies to Evaluate Statistical Methods.” Statistics in Medicine, ahead of print, January. https://doi.org/10.1002/sim.8086.
Oehlert, Gary W. 1992. “A Note on the Delta Method.” The American Statistician 46 (1): 27–29. https://doi.org/10.1080/00031305.1992.10475842.
Pashley, Nicole E, Luke Keele, and Luke W Miratrix. 2024. “Improving Instrumental Variable Estimators with Poststratification.” Journal of the Royal Statistical Society Series A: Statistics in Society, qnae073.
Pedersen, Thomas Lin. 2024. Patchwork: The Composer of Plots. https://patchwork.data-imaginist.com.
Pustejovsky, James E. 2014. “Converting from d to r to z When the Design Uses Extreme Groups, Dichotomization, or Experimental Control.” Psychological Methods 19 (1): 92.
Pustejovsky, James E. 2018. The Multivariate Delta Method. https://jepusto.com/posts/Multivariate-delta-method/.
Robert, Christian, and George Casella. 2010. Introducing Monte Carlo Methods with R. Springer. https://doi.org/10.1007/978-1-4419-1576-4.
Rosenbaum, Paul R. 2017. Observation and Experiment: An Introduction to Causal Inference. Harvard University Press.
Rousseeuw, Peter J., and Christophe Croux. 1993. “Alternatives to the Median Absolute Deviation.” Journal of the American Statistical Association 88 (424): 1273–83. https://doi.org/10.1080/01621459.1993.10476408.
Satterthwaite, F. E. 1946. “An Approximate Distribution of Estimates of Variance Components.” Biometrics Bulletin 2 (6): 110. https://doi.org/10.2307/3002019.
Shaw, Pamela A., Susan Gruber, Brian D. Williamson, et al. 2025. A Cautionary Note for Plasmode Simulation Studies in the Setting of Causal Inference. https://arxiv.org/abs/2504.11740.
Siepe, Björn S., František Bartoš, Tim P. Morris, Anne-Laure Boulesteix, Daniel W. Heck, and Samuel Pawel. 2024. “Simulation Studies for Methodological Research in Psychology: A Standardized Template for Planning, Preregistration, and Reporting.” Psychological Methods, ahead of print. https://doi.org/10.1037/met0000695.
Staiger, Douglas O, and Jonah E Rockoff. 2010. “Searching for Effective Teachers with Imperfect Information.” Journal of Economic Perspectives 24 (3): 97–118.
Tipton, Elizabeth. 2013. “Stratified Sampling Using Cluster Analysis: A Sample Selection Strategy for Improved Generalizations from Experiments.” Evaluation Review 37 (2): 109–39. https://doi.org/10.1177/0193841X13516324.
Tipton, Elizabeth, and James E Pustejovsky. 2015. “Small-Sample Adjustments for Tests of Moderators and Model Fit Using Robust Variance Estimation in Meta-Regression.” Journal of Educational and Behavioral Statistics 40 (6): 604–34.
Tufte, Edward R, and Peter R Graves-Morris. 1983. The Visual Display of Quantitative Information. Graphics press Cheshire, CT.
Vevea, Jack L, and Larry V Hedges. 1995. “A General Linear Model for Estimating Effect Size in the Presence of Publication Bias.” Psychometrika 60 (3): 419–35. https://doi.org/10.1007/BF02294384.
Welch, B. L. 1951. “On the Comparison of Several Mean Values: An Alternative Approach.” Biometrika 38 (3/4): 330. https://doi.org/10.2307/2332579.
Westfall, Peter H, and Kevin SS Henning. 2013. Understanding Advanced Statistical Methods. Vol. 543. CRC Press Boca Raton, FL.
White, Halbert. 1980. “A Heteroskedasticity-Consistent Covariance Matrix Estimator and a Direct Test for Heteroskedasticity.” Econometrica 48 (4): 817–38.
Wickham, Hadley. 2014. “Tidy Data.” Journal of Statistical Software 59 (10): 1–23. https://doi.org/10.18637/jss.v059.i10.
Wickham, Hadley, and Jennifer Bryan. 2023. R Packages. "O’Reilly Media, Inc." https://r-pkgs.org/.
Wickham, Hadley, Mine Çetinkaya-Rundel, and Garrett Grolemund. 2023. R for Data Science: Import, Tidy, Transform, Visualize, and Model Data. "O’Reilly Media, Inc." https://r4ds.hadley.nz/.
Wilcox, Rand R. 2022. Introduction to Robust Estimation and Hypothesis Testing. Fifth edition. Academic Press, an imprint of Elsevier.
Zhou, Jie, Zhiwei Zhang, Zhaohai Li, and Jun Zhang. 2015. “Coarsened Propensity Scores and Hybrid Estimators for Missing Data and Causal Inference.” International Statistical Review 83 (3): 449–71.