Skip to content

od p-value computation disagrees with od/pytorch/ensemble.py's own denominator #958

Description

@shivamlalakiya

Summary

od/pytorch/ensemble.py:114 computes an outlier p-value as

p_vals = (1 + less_than_val_scores.sum(1))/(len(self.val_scores) + 1)

od/pytorch/base.py:199 and od/sklearn/base.py:148 compute the same
quantity in the non-ensemble path, one line each, but with a different
denominator and a different comparator:

return (1 + (scores[:, None] < self.val_scores).sum(-1))/len(self.val_scores) \
    if self.threshold_inferred else None

Same author, same subpackage, same "add one to the count of validation
scores the test score beats" idea, two different denominators
(len(self.val_scores) vs. len(self.val_scores) + 1) and two different
comparators (< vs. presumably <= for the ensemble version's
less_than_val_scores, computed elsewhere). Wanted to flag the
inconsistency rather than guess which one is intended.

What the base.py denominator produces

With len(self.val_scores) in the denominator, a test score below every
validation score returns (n+1)/n, which is greater than 1. Reproduced
through the public GMM detector, 0.13.0 wheel from PyPI, Python 3.11,
numpy 1.26.4:

>>> import numpy as np
>>> from alibi_detect.od._gmm import GMM
>>> X_fit = np.array([[0.0],[1.0],[2.0],[3.0],[4.0],[5.0],[6.0],[7.0]], dtype=np.float32)
>>> X_cal = np.array([[1.0],[2.0],[3.0],[4.0]], dtype=np.float32)
>>> d = GMM(n_components=1, backend='sklearn')
>>> d.fit(X_fit)
>>> d.infer_threshold(X_cal, fpr=0.5)
>>> d.predict(np.array([[3.5]], dtype=np.float32))['data']['p_value']
array([1.25])

len(val_scores) here is 4, and (4+1)/4 = 1.25, in the documented
p_value output field.

What the strict < produces

Calling _p_vals directly with validation scores tied to the test score
(9 and 19 calibration points, all equal to the test score):

>>> s.val_scores = np.full(9, 5.0); s._p_vals(np.array([5.0]))
array([0.11111111])   # 1/9
>>> s.val_scores = np.full(19, 5.0); s._p_vals(np.array([5.0]))
array([0.05263158])   # 1/19

Exact-rational enumeration over the attained support (sup_t[P(p<=t) - t],
computed with fractions.Fraction, no Monte Carlo) gives, tie-free vs.
fully tied vs. two equal-size blocks:

n= 9   tie-free 0.0000   all tied 0.8889   two blocks 0.3889
n=19   tie-free 0.0000   all tied 0.9474   two blocks 0.4474

Direction

These are two separate effects pulling in opposite directions: the
denominator makes the value too large (up to impossible, >1), and the
strict comparator under ties makes it too small. They don't cancel in
general, since they depend on different properties of the input (how
close the test score is to the extreme of the validation set, vs. how
many exact ties there are).

PR #959 changes both lines to <= over len(self.val_scores) + 1,
matching ensemble.py's denominator. ensemble.py:114 itself still uses
the strict comparator via less_than_val_scores computed above it, so
the tie side of this is package-wide; I didn't touch ensemble.py in the
PR since its denominator is already correct and I didn't want to bundle
an unrelated comparator change into that file.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions