As an extra piece of analysis to demonstrate where this approach fits within wider work:
We could generate a comparison of model rankings using each evaluation study design (in the table of approaches to increasingly formal designs), ie none; restriction; stratification; comparison to baseline, etc; and compare how model rankings change with each method.
Specifically this suggestion started with the idea of comparing each model's ranking between: the standard pairwise-relative-skill method we've used elsewhere; to models ranked by the "unexplained" variance left after controlling for the set of target-related confounding factors (but not controlling for any model specification factors, unlike in the main study here).
This would help show the difference made by increasingly controlled/nuanced evaluation methods.
Could add this directly or as supplementary analysis.
Or leave for future work.
(Ideally this is the kind of thing we would be able to reproduce and extend some piece of existing evaluation, i.e. of FluSight.)
Thanks to Gneiting group with @sbfnk for suggestion.
As an extra piece of analysis to demonstrate where this approach fits within wider work:
We could generate a comparison of model rankings using each evaluation study design (in the table of approaches to increasingly formal designs), ie none; restriction; stratification; comparison to baseline, etc; and compare how model rankings change with each method.
Specifically this suggestion started with the idea of comparing each model's ranking between: the standard pairwise-relative-skill method we've used elsewhere; to models ranked by the "unexplained" variance left after controlling for the set of target-related confounding factors (but not controlling for any model specification factors, unlike in the main study here).
This would help show the difference made by increasingly controlled/nuanced evaluation methods.
Could add this directly or as supplementary analysis.
Or leave for future work.
(Ideally this is the kind of thing we would be able to reproduce and extend some piece of existing evaluation, i.e. of FluSight.)
Thanks to Gneiting group with @sbfnk for suggestion.