Skip to content
Edit this page View source of this page

API Reference

Reference documentation for the complete v1 public module surface. For the statistical assumptions and decision context behind an API, follow the linked user-guide page before relying on a guarantee.

Start Here

If you are looking for task-oriented call sequences, start with Common Workflows.

For the v1 public compatibility contract, see API Stability.

Detector

nonconform.detector

Core conformal anomaly detector implementation.

This module provides :class:ConformalDetector, which calibrates scores from a supported anomaly detector. The default empirical estimator returns rank-based conformal p-values; other estimators document their own interpretation. Weighted mode estimates density-ratio weights for covariate-shift workflows. Selection methods apply a multiple-testing procedure to a complete test batch. The resulting validity and false discovery rate guarantees depend on the assumptions documented for the chosen calibration and selection procedure.

Classes:

Name Description
BaseConformalDetector

Abstract base class for conformal detectors.

ConformalDetector

Main conformal anomaly detector with optional weighting.

BaseConformalDetector

Bases: ABC

Abstract base class for all conformal anomaly detectors.

Defines the core interface that all conformal anomaly detection implementations must provide. Conformal detectors support either an integrated or detached calibration workflow:

  1. Integrated calibration: fit() trains detector(s) and computes calibration scores
  2. Detached calibration: train detector externally, then call calibrate() on a separate calibration dataset
  3. Inference phase: compute_p_values() applies the configured estimation strategy, while select() combines estimation with the configured batch selection procedure

Subclasses must implement both abstract methods.

Note

This is an abstract class and cannot be instantiated directly. Use ConformalDetector for the main implementation.

fit abstractmethod
fit(
    x: DataFrame | ndarray,
    y: ndarray | None = None,
    *,
    n_jobs: int | None = None,
) -> Self

Fit the detector model(s) and compute calibration scores.

Parameters:

Name Type Description Default
x DataFrame | ndarray

The dataset used for fitting the model(s) and determining calibration scores.

required
y ndarray | None

Ignored. Present for sklearn API compatibility.

None
n_jobs int | None

Optional strategy-specific parallelism hint. Currently used by strategies that expose an n_jobs parameter (for example, JackknifeBootstrap).

None

Returns:

Type Description
Self

The fitted detector instance.

Source code in nonconform/detector.py
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
@ensure_numpy_array
@abstractmethod
def fit(
    self,
    x: pd.DataFrame | np.ndarray,
    y: np.ndarray | None = None,
    *,
    n_jobs: int | None = None,
) -> Self:
    """Fit the detector model(s) and compute calibration scores.

    Args:
        x: The dataset used for fitting the model(s) and determining
            calibration scores.
        y: Ignored. Present for sklearn API compatibility.
        n_jobs: Optional strategy-specific parallelism hint.
            Currently used by strategies that expose an ``n_jobs`` parameter
            (for example, ``JackknifeBootstrap``).

    Returns:
        The fitted detector instance.
    """
    raise NotImplementedError("Subclasses must implement fit()")
calibrate
calibrate(
    x: DataFrame | ndarray, y: ndarray | None = None
) -> Self

Calibrate a pre-fitted detector on separate calibration data.

Parameters:

Name Type Description Default
x DataFrame | ndarray

Dataset used only to compute calibration scores.

required
y ndarray | None

Ignored. Present for sklearn API compatibility.

None

Returns:

Type Description
Self

The calibrated detector instance.

Source code in nonconform/detector.py
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
@ensure_numpy_array
def calibrate(
    self,
    x: pd.DataFrame | np.ndarray,
    y: np.ndarray | None = None,
) -> Self:
    """Calibrate a pre-fitted detector on separate calibration data.

    Args:
        x: Dataset used only to compute calibration scores.
        y: Ignored. Present for sklearn API compatibility.

    Returns:
        The calibrated detector instance.
    """
    raise NotImplementedError("Subclasses must implement calibrate()")
compute_p_values abstractmethod
compute_p_values(
    x: DataFrame | Series | ndarray,
    *,
    refit_weights: bool = True,
) -> np.ndarray | pd.Series

Return conformal p-values for new data.

Parameters:

Name Type Description Default
x DataFrame | Series | ndarray

New data instances for anomaly estimation.

required
refit_weights bool

Whether to refit the weight estimator for this batch in weighted mode. Ignored in standard mode.

True

Returns:

Type Description
ndarray | Series

P-values as ndarray for numpy input, or pandas Series for pandas input.

Source code in nonconform/detector.py
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
@abstractmethod
def compute_p_values(
    self,
    x: pd.DataFrame | pd.Series | np.ndarray,
    *,
    refit_weights: bool = True,
) -> np.ndarray | pd.Series:
    """Return conformal p-values for new data.

    Args:
        x: New data instances for anomaly estimation.
        refit_weights: Whether to refit the weight estimator for this batch
            in weighted mode. Ignored in standard mode.

    Returns:
        P-values as ndarray for numpy input, or pandas Series for pandas input.
    """
    raise NotImplementedError("Subclasses must implement compute_p_values()")
score_samples abstractmethod
score_samples(
    x: DataFrame | Series | ndarray,
    *,
    refit_weights: bool = True,
) -> np.ndarray | pd.Series

Return aggregated raw anomaly scores for new data.

Parameters:

Name Type Description Default
x DataFrame | Series | ndarray

New data instances for anomaly estimation.

required
refit_weights bool

Whether to refit the weight estimator for this batch in weighted mode. Ignored in standard mode.

True

Returns:

Type Description
ndarray | Series

Raw scores as ndarray for numpy input, or pandas Series for pandas input.

Source code in nonconform/detector.py
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
@abstractmethod
def score_samples(
    self,
    x: pd.DataFrame | pd.Series | np.ndarray,
    *,
    refit_weights: bool = True,
) -> np.ndarray | pd.Series:
    """Return aggregated raw anomaly scores for new data.

    Args:
        x: New data instances for anomaly estimation.
        refit_weights: Whether to refit the weight estimator for this batch
            in weighted mode. Ignored in standard mode.

    Returns:
        Raw scores as ndarray for numpy input, or pandas Series for pandas input.
    """
    raise NotImplementedError("Subclasses must implement score_samples()")

ConformalDetector

ConformalDetector(
    detector: Any,
    strategy: BaseStrategy,
    estimation: BaseEstimation | None = None,
    weight_estimator: BaseWeightEstimator | None = None,
    aggregation: str = "median",
    score_polarity: ScorePolarity
    | Literal[
        "auto", "higher_is_anomalous", "higher_is_normal"
    ]
    | None = None,
    seed: int | None = None,
    verbose: bool = False,
    verify_prepared_batch_content: bool = True,
)

Bases: BaseConformalDetector

Wrap an anomaly detector with conformal calibration and batch selection.

The wrapped detector may be a recognized scikit-learn estimator, a PyOD model, or a custom object implementing the :class:~nonconform.structures.AnomalyDetector protocol.

In standard mode, the fitted strategy supplies one or more fixed scoring rules and calibration-score sets. With the default Empirical estimator, compute_p_values() ranks each test score against those calibration scores. The usual marginal p-value validity statement requires exchangeability of the relevant null calibration and test examples. select() then applies Benjamini-Hochberg to the full test family. Any FDR guarantee additionally depends on the dependence assumptions of that multiple-testing procedure. Other estimation strategies document their own interpretation and assumptions.

Supplying weight_estimator enables weighted mode. The estimator learns density-ratio weights from calibration and test covariates, and select() uses weighted conformalized selection. Its validity requires the applicable covariate-shift assumptions, support overlap, and adequate weights; the class cannot verify those scientific assumptions from data alone.

With DerandomizedSplits, fit() retains repeated model/calibration pairs and select() uniformly aggregates per-split e-values before e-BH. Inspect last_selection_result for evidence and selection diagnostics. P-value methods and detached calibration are unavailable for this strategy.

Parameters:

Name Type Description Default
detector Any

Anomaly detector (PyOD, sklearn-compatible, or custom).

required
strategy BaseStrategy

The conformal strategy for fitting, calibration, and evidence construction. DerandomizedSplits selects through e-values and e-BH.

required
estimation BaseEstimation | None

P-value estimation strategy. Defaults to Empirical(). Unused by DerandomizedSplits, which accepts only None or ordinary Empirical.

None
weight_estimator BaseWeightEstimator | None

Weight estimator for covariate shift. Defaults to None.

None
aggregation str

Method for aggregating scores from multiple fitted models: "mean", "median", "minimum", or "maximum". Defaults to "median". For DerandomizedSplits, affects score_samples() only; selection always uniformly averages per-split e-values.

'median'
score_polarity ScorePolarity | Literal['auto', 'higher_is_anomalous', 'higher_is_normal'] | None

Score direction convention. Use "higher_is_anomalous" when higher raw scores indicate more anomalous samples, and "higher_is_normal" when higher scores indicate more normal samples. If omitted (None), nonconform applies an implicit default policy: known sklearn normality detectors resolve to "higher_is_normal", while PyOD and unknown custom detectors resolve to "higher_is_anomalous". Explicit "auto" enables strict inference: known detector families are inferred, and unknown detectors raise. Defaults to None.

None
seed int | None

Random seed for reproducibility. Defaults to None.

None
verbose bool

If True, displays aggregation progress for multi-model strategies. Defaults to False.

False
verify_prepared_batch_content bool

If True (default), weighted reuse mode (refit_weights=False) verifies exact batch content identity via hashing. This adds O(n) overhead per checked batch. Set to False to skip content hashing and validate only batch size.

True

Attributes:

Name Type Description
detector

The underlying anomaly detection model.

strategy

The calibration and evidence-construction strategy.

weight_estimator

Optional weight estimator for handling covariate shift.

aggregation

Method for combining scores from multiple models.

score_polarity ScorePolarity

Resolved score polarity used internally.

seed ScorePolarity

Random seed for reproducible results.

verbose ScorePolarity

Whether to display progress bars.

Examples:

Standard conformal p-values and batch selection:

import numpy as np
from sklearn.ensemble import IsolationForest

from nonconform import ConformalDetector, Split

rng = np.random.default_rng(42)
x_reference = rng.normal(size=(300, 2))
x_test = np.vstack([rng.normal(size=(38, 2)), rng.normal(loc=5.0, size=(2, 2))])
detector = ConformalDetector(
    detector=IsolationForest(random_state=42),
    strategy=Split(n_calib=0.25),
    score_polarity="higher_is_normal",
    seed=42,
)
detector.fit(x_reference)
p_values = detector.compute_p_values(x_test)
selected = detector.select(x_test, alpha=0.10)
print(p_values.shape, np.flatnonzero(selected))

Weighted conformal p-values under a simulated covariate shift:

import numpy as np
from sklearn.ensemble import IsolationForest

from nonconform import (
    ConformalDetector,
    Split,
    logistic_weight_estimator,
)

rng = np.random.default_rng(7)
x_reference = rng.normal(size=(400, 2))
x_test = rng.normal(loc=0.5, size=(40, 2))
detector = ConformalDetector(
    detector=IsolationForest(random_state=7),
    strategy=Split(n_calib=0.25),
    weight_estimator=logistic_weight_estimator(),
    score_polarity="higher_is_normal",
    seed=7,
)
detector.fit(x_reference)
p_values = detector.compute_p_values(x_test)
print(p_values.shape, p_values.min(), p_values.max())

Detached calibration with a pre-trained model (Split strategy):

import numpy as np
from sklearn.ensemble import IsolationForest

from nonconform import ConformalDetector, Split

rng = np.random.default_rng(11)
x_fit = rng.normal(size=(200, 2))
x_calibration = rng.normal(size=(100, 2))
x_test = rng.normal(size=(10, 2))
base_detector = IsolationForest(random_state=11).fit(x_fit)
detector = ConformalDetector(
    detector=base_detector,
    strategy=Split(n_calib=0.2),
    score_polarity="higher_is_normal",
)
detector.calibrate(x_calibration)
p_values = detector.compute_p_values(x_test)
print(p_values)
Note

Strict inductive conformal workflows require a fixed training-only score map at inference time. PyOD detectors known to violate this are: CD, COF, COPOD, ECOD, LMDD, LOCI, RGraph, SOD, SOS.

Source code in nonconform/detector.py
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
def __init__(
    self,
    detector: Any,
    strategy: BaseStrategy,
    estimation: BaseEstimation | None = None,
    weight_estimator: BaseWeightEstimator | None = None,
    aggregation: str = "median",
    score_polarity: ScorePolarity
    | Literal["auto", "higher_is_anomalous", "higher_is_normal"]
    | None = None,
    seed: int | None = None,
    verbose: bool = False,
    verify_prepared_batch_content: bool = True,
) -> None:
    self._configure(
        detector=detector,
        strategy=strategy,
        estimation=estimation,
        weight_estimator=weight_estimator,
        aggregation=aggregation,
        score_polarity=score_polarity,
        seed=seed,
        verbose=verbose,
        verify_prepared_batch_content=verify_prepared_batch_content,
    )
detector_set property
detector_set: list[AnomalyDetector]

Returns a copy of the list of trained detector models.

calibration_set property
calibration_set: ndarray

Return a copy of calibration scores.

DerandomizedSplits returns a matrix shaped (n_repetitions, n_calibration), with one row per retained model. Other strategies return their existing one-dimensional score array.

calibration_samples property
calibration_samples: ndarray

Returns a copy of the calibration samples (weighted mode only).

last_result property
last_result: ConformalResult | None

Return the most recent raw-score or p-value snapshot.

DerandomizedSplits selection instead populates last_selection_result and clears this snapshot. Its raw-score snapshots have no pooled calibration scores.

last_selection_result property
last_selection_result: EValueSelectionResult | None

Return a defensive snapshot of the latest e-value selection.

None before selection, after fitting or raw scoring, and for existing p-value selection workflows. Arrays remain read-only in the snapshot.

score_polarity property
score_polarity: ScorePolarity

Returns the resolved score polarity convention.

is_fitted property
is_fitted: bool

Returns whether the detector has been fitted.

get_params
get_params(deep: bool = True) -> dict[str, Any]

Return estimator parameters following sklearn conventions.

Notes
  • deep=False returns constructor-facing parameters used for sklearn clone compatibility.
  • deep=True also includes nested component__param entries read from the current runtime components (effective/internal state), which may differ from originally passed constructor objects after adaptation/normalization.
Source code in nonconform/detector.py
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
def get_params(self, deep: bool = True) -> dict[str, Any]:
    """Return estimator parameters following sklearn conventions.

    Notes:
        - ``deep=False`` returns constructor-facing parameters used for
          sklearn clone compatibility.
        - ``deep=True`` also includes nested ``component__param`` entries
          read from the current runtime components (effective/internal state),
          which may differ from originally passed constructor objects after
          adaptation/normalization.
    """
    params: dict[str, Any] = {
        "detector": self._init_detector,
        "strategy": self._init_strategy,
        "estimation": self._init_estimation,
        "weight_estimator": self._init_weight_estimator,
        "aggregation": self._init_aggregation,
        "score_polarity": self._init_score_polarity,
        "seed": self._init_seed,
        "verbose": self._init_verbose,
        "verify_prepared_batch_content": self._init_verify_prepared_batch_content,
    }
    if not deep:
        return params

    for component_name in self._NESTED_COMPONENTS:
        component = getattr(self, component_name)
        if component is None or not hasattr(component, "get_params"):
            continue
        try:
            component_params = component.get_params(deep=True)
        except TypeError:
            component_params = component.get_params()
        for key, value in component_params.items():
            params[f"{component_name}__{key}"] = value
    return params
set_params
set_params(**params: Any) -> Self

Set estimator parameters following sklearn conventions.

Source code in nonconform/detector.py
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
def set_params(self, **params: Any) -> Self:
    """Set estimator parameters following sklearn conventions."""
    if not params:
        return self

    updated_params = self.get_params(deep=False)
    nested_updates: dict[str, dict[str, Any]] = {}

    for key, value in params.items():
        if "__" in key:
            component_name, nested_key = key.split("__", 1)
            if component_name not in self._NESTED_COMPONENTS:
                raise ValueError(f"Invalid parameter {component_name!r}.")
            nested_updates.setdefault(component_name, {})[nested_key] = value
            continue

        if key not in updated_params:
            raise ValueError(
                f"Invalid parameter {key!r} for estimator {type(self).__name__}."
            )
        updated_params[key] = value

    for component_name, component_params in nested_updates.items():
        component = updated_params[component_name]
        if component is None:
            raise ValueError(
                f"Cannot set nested parameters for {component_name!r}: "
                "component is None."
            )
        if not hasattr(component, "set_params"):
            raise ValueError(
                f"Cannot set nested parameters for {component_name!r}: "
                "component does not implement set_params()."
            )
        component.set_params(**component_params)

    self._configure(**updated_params)
    return self
fit
fit(
    x: DataFrame | ndarray,
    y: ndarray | None = None,
    *,
    n_jobs: int | None = None,
) -> Self

Fit detector model(s) and compute calibration scores.

Uses the specified strategy to train the base detector(s) and calculate non-conformity scores on the calibration set.

Parameters:

Name Type Description Default
x DataFrame | ndarray

The dataset used for fitting and calibration.

required
y ndarray | None

Ignored. Present for sklearn API compatibility.

None
n_jobs int | None

Optional strategy-specific parallelism hint. Supported by strategies whose fit_calibrate signature includes n_jobs (for example, JackknifeBootstrap).

None

Returns:

Type Description
Self

The fitted detector instance (for method chaining).

Source code in nonconform/detector.py
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
@ensure_numpy_array
def fit(
    self,
    x: pd.DataFrame | np.ndarray,
    y: np.ndarray | None = None,
    *,
    n_jobs: int | None = None,
) -> Self:
    """Fit detector model(s) and compute calibration scores.

    Uses the specified strategy to train the base detector(s) and calculate
    non-conformity scores on the calibration set.

    Args:
        x: The dataset used for fitting and calibration.
        y: Ignored. Present for sklearn API compatibility.
        n_jobs: Optional strategy-specific parallelism hint. Supported by
            strategies whose ``fit_calibrate`` signature includes ``n_jobs``
            (for example, ``JackknifeBootstrap``).

    Returns:
        The fitted detector instance (for method chaining).
    """
    _ = y
    self._last_selection_result = None
    if self.strategy._uses_e_values:
        self._reset_fit_state()
    fit_kwargs: dict[str, Any] = {
        "x": x,
        "detector": self.detector,
        "weighted": self._is_weighted_mode,
        "seed": self.seed,
    }
    if n_jobs is not None:
        strategy_params = inspect.signature(self.strategy.fit_calibrate).parameters
        if "n_jobs" not in strategy_params:
            raise ValueError(
                f"Strategy {type(self.strategy).__name__} does not support n_jobs. "
                "Pass n_jobs only when using a strategy that exposes it, "
                "such as JackknifeBootstrap."
            )
        fit_kwargs["n_jobs"] = n_jobs

    self._detector_set, self._calibration_set = self.strategy.fit_calibrate(
        **fit_kwargs
    )
    self._calibration_mode = CalibrationMode.INTEGRATED
    self._n_features_in = int(x.shape[1])

    if (
        self._is_weighted_mode
        and self.strategy.calibration_ids is not None
        and len(self.strategy.calibration_ids) > 0
    ):
        self._calibration_samples = x[self.strategy.calibration_ids]
    else:
        self._calibration_samples = np.array([])

    self._prepared_weight_batch_size = None
    self._prepared_weight_batch_signature = None
    self._last_result = None
    return self
calibrate
calibrate(
    x: DataFrame | ndarray, y: ndarray | None = None
) -> Self

Calibrate a pre-fitted detector on separate calibration data.

This detached workflow is currently supported only for Split strategy, where a single pre-fitted model is calibrated on a dedicated dataset.

Parameters:

Name Type Description Default
x DataFrame | ndarray

Calibration dataset used to compute calibration scores.

required
y ndarray | None

Ignored. Present for sklearn API compatibility.

None

Returns:

Type Description
Self

The calibrated detector instance (for method chaining).

Raises:

Type Description
ValueError

If strategy is not Split.

NotFittedError

If the base detector appears unfitted.

Source code in nonconform/detector.py
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
@ensure_numpy_array
def calibrate(
    self,
    x: pd.DataFrame | np.ndarray,
    y: np.ndarray | None = None,
) -> Self:
    """Calibrate a pre-fitted detector on separate calibration data.

    This detached workflow is currently supported only for ``Split`` strategy,
    where a single pre-fitted model is calibrated on a dedicated dataset.

    Args:
        x: Calibration dataset used to compute calibration scores.
        y: Ignored. Present for sklearn API compatibility.

    Returns:
        The calibrated detector instance (for method chaining).

    Raises:
        ValueError: If strategy is not ``Split``.
        NotFittedError: If the base detector appears unfitted.
    """
    _ = y
    self._last_selection_result = None
    if not isinstance(self.strategy, Split):
        raise ValueError(
            "calibrate() is supported only with Split strategy. "
            f"Got {type(self.strategy).__name__}. Use fit(x_reference) for "
            "integrated calibration."
        )

    try:
        calibration_set = np.asarray(
            self.detector.decision_function(x),
            dtype=float,
        ).ravel()
    except Exception as exc:
        message = str(exc).lower()
        if (
            isinstance(exc, NotFittedError)
            or "not fitted" in message
            or (isinstance(exc, AttributeError) and "has no attribute" in message)
        ):
            raise NotFittedError(
                "Base detector is not fitted. Fit the base detector before "
                "calling calibrate()."
            ) from exc
        raise

    if calibration_set.shape[0] != len(x):
        raise ValueError(
            "calibration scores must have one value per calibration sample. "
            f"Got {calibration_set.shape[0]} scores for {len(x)} samples."
        )

    self._detector_set = [self.detector]
    self._calibration_set = calibration_set
    self._calibration_mode = CalibrationMode.DETACHED
    self._n_features_in = int(x.shape[1])
    if self._is_weighted_mode:
        self._calibration_samples = x.copy()
    else:
        self._calibration_samples = np.array([])

    self._prepared_weight_batch_size = None
    self._prepared_weight_batch_signature = None
    self._last_result = None
    return self
select
select(
    x: DataFrame | Series | ndarray,
    *,
    alpha: float = 0.05,
    pruning: Pruning = Pruning.DETERMINISTIC,
    seed: int | None = None,
    refit_weights: bool = True,
) -> np.ndarray | pd.Series

Construct evidence and select anomalies from one fixed test batch.

This is the single-call batch workflow. It combines compute_p_values() with Benjamini-Hochberg in standard mode or weighted conformalized selection in weighted mode. Validity still depends on the assumptions of both the p-value construction and the selected multiple-testing procedure.

With DerandomizedSplits, this instead constructs per-split e-values, averages them uniformly, and applies e-BH once. Configure alpha_bh and tie_seed on that strategy. Automatic ties use a separate fitting-derived seed. Selection populates last_selection_result and clears last_result.

Parameters:

Name Type Description Default
x DataFrame | Series | ndarray

New data instances for anomaly estimation.

required
alpha float

Nominal FDR target in (0, 1). Defaults to 0.05.

0.05
pruning Pruning

Pruning strategy for weighted FDR control. Ignored in standard (unweighted) mode. Defaults to Pruning.DETERMINISTIC.

DETERMINISTIC
seed int | None

Optional random seed for weighted randomized pruning modes. When None, falls back to detector seed. Ignored in standard mode and deterministic pruning mode.

None
refit_weights bool

Whether to refit the weight estimator for this batch in weighted mode. Ignored in standard mode. Defaults to True.

True

Returns:

Type Description
ndarray | Series

Boolean selection mask of shape (n_test,). True entries are

ndarray | Series

the selected anomaly discoveries. Returns a pandas Series when the

ndarray | Series

input is a DataFrame or Series.

Examples:

Standard workflow (no weight estimator):

import numpy as np
from sklearn.ensemble import IsolationForest

from nonconform import ConformalDetector, Split

rng = np.random.default_rng(42)
x_reference = rng.normal(size=(300, 2))
x_test = np.vstack(
    [rng.normal(size=(38, 2)), rng.normal(loc=5.0, size=(2, 2))]
)
detector = ConformalDetector(
    detector=IsolationForest(random_state=42),
    strategy=Split(n_calib=0.25),
    score_polarity="higher_is_normal",
    seed=42,
).fit(x_reference)
selected = detector.select(x_test, alpha=0.10)
print("Selected indices:", np.flatnonzero(selected))

Weighted workflow:

import numpy as np
from sklearn.ensemble import IsolationForest

from nonconform import (
    ConformalDetector,
    Split,
    logistic_weight_estimator,
)

rng = np.random.default_rng(7)
x_reference = rng.normal(size=(400, 2))
x_test = rng.normal(loc=0.5, size=(40, 2))
detector = ConformalDetector(
    detector=IsolationForest(random_state=7),
    strategy=Split(n_calib=0.25),
    weight_estimator=logistic_weight_estimator(),
    score_polarity="higher_is_normal",
    seed=7,
).fit(x_reference)
selected = detector.select(x_test, alpha=0.10)
print("Number selected:", int(selected.sum()))
Source code in nonconform/detector.py
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
939
940
941
942
943
944
def select(
    self,
    x: pd.DataFrame | pd.Series | np.ndarray,
    *,
    alpha: float = 0.05,
    pruning: Pruning = Pruning.DETERMINISTIC,
    seed: int | None = None,
    refit_weights: bool = True,
) -> np.ndarray | pd.Series:
    """Construct evidence and select anomalies from one fixed test batch.

    This is the single-call batch workflow. It combines
    ``compute_p_values()`` with Benjamini-Hochberg in standard mode or
    weighted conformalized selection in weighted mode. Validity still
    depends on the assumptions of both the p-value construction and the
    selected multiple-testing procedure.

    With DerandomizedSplits, this instead constructs per-split e-values,
    averages them uniformly, and applies e-BH once. Configure alpha_bh and
    tie_seed on that strategy. Automatic ties use a separate fitting-derived
    seed. Selection populates last_selection_result and clears last_result.

    Args:
        x: New data instances for anomaly estimation.
        alpha: Nominal FDR target in ``(0, 1)``. Defaults to ``0.05``.
        pruning: Pruning strategy for weighted FDR control. Ignored in
            standard (unweighted) mode. Defaults to
            ``Pruning.DETERMINISTIC``.
        seed: Optional random seed for weighted randomized pruning modes.
            When ``None``, falls back to detector ``seed``. Ignored in
            standard mode and deterministic pruning mode.
        refit_weights: Whether to refit the weight estimator for this batch
            in weighted mode. Ignored in standard mode. Defaults to True.

    Returns:
        Boolean selection mask of shape ``(n_test,)``. ``True`` entries are
        the selected anomaly discoveries. Returns a pandas Series when the
        input is a DataFrame or Series.

    Examples:
        Standard workflow (no weight estimator):

        ```python
        import numpy as np
        from sklearn.ensemble import IsolationForest

        from nonconform import ConformalDetector, Split

        rng = np.random.default_rng(42)
        x_reference = rng.normal(size=(300, 2))
        x_test = np.vstack(
            [rng.normal(size=(38, 2)), rng.normal(loc=5.0, size=(2, 2))]
        )
        detector = ConformalDetector(
            detector=IsolationForest(random_state=42),
            strategy=Split(n_calib=0.25),
            score_polarity="higher_is_normal",
            seed=42,
        ).fit(x_reference)
        selected = detector.select(x_test, alpha=0.10)
        print("Selected indices:", np.flatnonzero(selected))
        ```

        Weighted workflow:

        ```python
        import numpy as np
        from sklearn.ensemble import IsolationForest

        from nonconform import (
            ConformalDetector,
            Split,
            logistic_weight_estimator,
        )

        rng = np.random.default_rng(7)
        x_reference = rng.normal(size=(400, 2))
        x_test = rng.normal(loc=0.5, size=(40, 2))
        detector = ConformalDetector(
            detector=IsolationForest(random_state=7),
            strategy=Split(n_calib=0.25),
            weight_estimator=logistic_weight_estimator(),
            score_polarity="higher_is_normal",
            seed=7,
        ).fit(x_reference)
        selected = detector.select(x_test, alpha=0.10)
        print("Number selected:", int(selected.sum()))
        ```
    """
    self._last_selection_result = None
    if self.strategy._uses_e_values:
        self._last_result = None
    if not (0.0 < alpha < 1.0):
        raise ValueError(f"alpha must be in (0, 1), got {alpha}")

    from nonconform.fdr import weighted_false_discovery_control

    x_array, index = _as_numpy_with_index(x)
    if self.strategy._uses_e_values:
        scores = self._score_models(x_array)
        selection = self.strategy._select_e_values(
            scores, self._calibration_set, alpha=alpha
        )
        self._last_selection_result = selection
        mask = selection.selected.copy()
        if index is not None:
            return pd.Series(mask, index=index, name="selected")
        return mask

    self.compute_p_values(x_array, refit_weights=refit_weights)
    result = self._last_result
    if result is None or result.p_values is None:
        raise RuntimeError(
            "Internal error: select() expected p-values after compute_p_values()."
        )

    if self._is_weighted_mode:
        selection_seed = self.seed if seed is None else seed
        mask = weighted_false_discovery_control(
            result=result,
            alpha=alpha,
            pruning=pruning,
            seed=_derive_wcs_pruning_seed(selection_seed),
        )
    else:
        p_values = np.asarray(result.p_values, dtype=float)
        mask = false_discovery_control(p_values, method="bh") <= alpha

    if index is not None:
        return pd.Series(mask, index=index, name="selected")
    return mask
prepare_weights_for
prepare_weights_for(x: DataFrame | ndarray) -> Self

Prepare weighted conformal state for a specific test batch.

In weighted mode, this fits the weight estimator for the supplied batch without producing predictions. Use this for explicit state transitions in exploratory workflows.

Parameters:

Name Type Description Default
x DataFrame | ndarray

Test batch for which weights should be prepared.

required

Returns:

Type Description
Self

The fitted detector instance (for method chaining).

Raises:

Type Description
NotFittedError

If fit() has not been called.

RuntimeError

If weighted mode is disabled.

Source code in nonconform/detector.py
946
947
948
949
950
951
952
953
954
955
956
957
958
959
960
961
962
963
964
965
966
967
968
969
970
971
972
@ensure_numpy_array
def prepare_weights_for(self, x: pd.DataFrame | np.ndarray) -> Self:
    """Prepare weighted conformal state for a specific test batch.

    In weighted mode, this fits the weight estimator for the supplied batch
    without producing predictions. Use this for explicit state transitions in
    exploratory workflows.

    Args:
        x: Test batch for which weights should be prepared.

    Returns:
        The fitted detector instance (for method chaining).

    Raises:
        NotFittedError: If fit() has not been called.
        RuntimeError: If weighted mode is disabled.
    """
    if not self.is_fitted:
        raise NotFittedError("This ConformalDetector instance is not fitted yet.")
    if not self._is_weighted_mode or self.weight_estimator is None:
        raise RuntimeError(
            "prepare_weights_for() requires weighted mode with a weight_estimator."
        )

    self._fit_weights_for_batch(x)
    return self
score_samples
score_samples(
    x: DataFrame | Series | ndarray,
    *,
    refit_weights: bool = True,
) -> np.ndarray | pd.Series

Return aggregated raw anomaly scores for new data.

Clears last_selection_result. With DerandomizedSplits, raw aggregation is diagnostic only and the last_result snapshot has calib_scores=None; select() instead uses each model's separate score/calibration pair.

Parameters:

Name Type Description Default
x DataFrame | Series | ndarray

New data instances for anomaly estimation.

required
refit_weights bool

Whether to refit the weight estimator for this batch in weighted mode. Defaults to True.

True

Returns:

Type Description
ndarray | Series

Aggregated raw anomaly scores.

Source code in nonconform/detector.py
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
def score_samples(
    self,
    x: pd.DataFrame | pd.Series | np.ndarray,
    *,
    refit_weights: bool = True,
) -> np.ndarray | pd.Series:
    """Return aggregated raw anomaly scores for new data.

    Clears last_selection_result. With DerandomizedSplits, raw aggregation
    is diagnostic only and the last_result snapshot has calib_scores=None;
    select() instead uses each model's separate score/calibration pair.

    Args:
        x: New data instances for anomaly estimation.
        refit_weights: Whether to refit the weight estimator for this batch
            in weighted mode. Defaults to True.

    Returns:
        Aggregated raw anomaly scores.
    """
    self._last_selection_result = None
    if self.strategy._uses_e_values:
        self._last_result = None
    x_array, index = _as_numpy_with_index(x)
    test_batch_signature = batch_signature(x_array)
    estimates = self._aggregate_scores(x_array)
    weights = self._resolve_weights(
        x_array,
        refit_weights=refit_weights,
        test_batch_signature=test_batch_signature,
    )
    calib_weights, test_weights = weights if weights else (None, None)

    result = ConformalResult(
        p_values=None,
        test_scores=estimates.copy(),
        calib_scores=(
            None if self.strategy._uses_e_values else self._calibration_set.copy()
        ),
        test_weights=_safe_copy(test_weights),
        calib_weights=_safe_copy(calib_weights),
        metadata={},
    )
    result._provenance = self._result_provenance(test_batch_signature)
    self._last_result = result
    if index is not None:
        return pd.Series(estimates, index=index, name="score")
    return estimates
compute_p_value
compute_p_value(x: Series | ndarray) -> float

Return one value from the configured estimation strategy.

Unavailable for DerandomizedSplits; use select() on a fixed test batch and inspect last_selection_result instead.

This is a single-sample convenience wrapper around :meth:compute_p_values. It updates :attr:last_result with the corresponding one-row result and does not update the calibration set. With the default Empirical estimator, the returned value is a rank-based conformal p-value. Other estimators define their own interpretation and assumptions.

Parameters:

Name Type Description Default
x Series | ndarray

One-dimensional feature vector with the same number of features used during fitting or detached calibration.

required

Returns:

Type Description
float

The observation's p-value or score-tail estimate as a Python float.

Raises:

Type Description
NotFittedError

If the detector has not been fitted or calibrated.

RuntimeError

If weighted conformal mode is enabled. Density-ratio estimation requires a representative test batch.

ValueError

If x is not one-dimensional or its feature count does not match the fitted data.

Note

Repeated single-sample calls are batch-equivalent only when detector scoring and p-value estimation are sample-wise and deterministic. Randomized tie-breaking can produce different values from one batch call.

Source code in nonconform/detector.py
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
def compute_p_value(self, x: pd.Series | np.ndarray) -> float:
    """Return one value from the configured estimation strategy.

    Unavailable for DerandomizedSplits; use select() on a fixed test batch
    and inspect last_selection_result instead.

    This is a single-sample convenience wrapper around
    :meth:`compute_p_values`. It updates :attr:`last_result` with the
    corresponding one-row result and does not update the calibration set.
    With the default ``Empirical`` estimator, the returned value is a
    rank-based conformal p-value. Other estimators define their own
    interpretation and assumptions.

    Args:
        x: One-dimensional feature vector with the same number of features
            used during fitting or detached calibration.

    Returns:
        The observation's p-value or score-tail estimate as a Python float.

    Raises:
        NotFittedError: If the detector has not been fitted or calibrated.
        RuntimeError: If weighted conformal mode is enabled. Density-ratio
            estimation requires a representative test batch.
        ValueError: If ``x`` is not one-dimensional or its feature count does
            not match the fitted data.

    Note:
        Repeated single-sample calls are batch-equivalent only when detector
        scoring and p-value estimation are sample-wise and deterministic.
        Randomized tie-breaking can produce different values from one batch
        call.
    """
    self._require_p_value_strategy()
    if not self.is_fitted:
        raise NotFittedError("This ConformalDetector instance is not fitted yet.")
    if self._is_weighted_mode:
        raise RuntimeError(
            "compute_p_value() is unavailable in weighted mode because "
            "density-ratio estimation requires a representative test batch. "
            "Use compute_p_values() with that batch instead."
        )

    x_array = np.asarray(x)
    if x_array.ndim != 1:
        raise ValueError(
            "x must be a one-dimensional feature vector; "
            f"got shape {x_array.shape}."
        )

    expected_features = self._n_features_in
    if expected_features is None:
        raise RuntimeError(
            "Fitted feature count is unavailable. Refit or recalibrate the "
            "detector before calling compute_p_value()."
        )
    if x_array.shape[0] != expected_features:
        raise ValueError(
            f"x has {x_array.shape[0]} features, but this ConformalDetector "
            f"was fitted with {expected_features} features."
        )

    p_values = self.compute_p_values(x_array[np.newaxis, :])
    return float(np.asarray(p_values).item())
compute_p_values
compute_p_values(
    x: DataFrame | Series | ndarray,
    *,
    refit_weights: bool = True,
) -> np.ndarray | pd.Series

Return values from the configured estimation strategy for new data.

Unavailable for DerandomizedSplits; use select() and inspect last_selection_result instead.

Parameters:

Name Type Description Default
x DataFrame | Series | ndarray

New data instances for anomaly estimation.

required
refit_weights bool

Whether to refit the weight estimator for this batch in weighted mode. Defaults to True.

True

Returns:

Type Description
ndarray | Series

P-values or score-tail estimates. Pandas input produces a Series

ndarray | Series

named "p_value"; NumPy input produces an ndarray.

Source code in nonconform/detector.py
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
def compute_p_values(
    self,
    x: pd.DataFrame | pd.Series | np.ndarray,
    *,
    refit_weights: bool = True,
) -> np.ndarray | pd.Series:
    """Return values from the configured estimation strategy for new data.

    Unavailable for DerandomizedSplits; use select() and inspect
    last_selection_result instead.

    Args:
        x: New data instances for anomaly estimation.
        refit_weights: Whether to refit the weight estimator for this batch
            in weighted mode. Defaults to True.

    Returns:
        P-values or score-tail estimates. Pandas input produces a Series
        named ``"p_value"``; NumPy input produces an ndarray.
    """
    self._require_p_value_strategy()
    x_array, index = _as_numpy_with_index(x)
    test_batch_signature = batch_signature(x_array)
    estimates = self._aggregate_scores(x_array)
    weights = self._resolve_weights(
        x_array,
        refit_weights=refit_weights,
        test_batch_signature=test_batch_signature,
    )
    calib_weights, test_weights = weights if weights else (None, None)

    p_values = self.estimation.compute_p_values(
        estimates, self._calibration_set, weights
    )

    metadata = self._result_metadata()
    if hasattr(self.estimation, "get_metadata"):
        meta = self.estimation.get_metadata()
        if meta:
            metadata.update(meta)

    result = ConformalResult(
        p_values=p_values.copy(),
        test_scores=estimates.copy(),
        calib_scores=self._calibration_set.copy(),
        test_weights=_safe_copy(test_weights),
        calib_weights=_safe_copy(calib_weights),
        metadata=metadata,
    )
    result._provenance = self._result_provenance(test_batch_signature)
    self._last_result = result
    if index is not None:
        return pd.Series(p_values, index=index, name="p_value")
    return p_values
fdp_bounds
fdp_bounds(
    x: DataFrame | Series | ndarray,
    *,
    confidence: float = 0.95,
    method: str = "mc_thc",
    n_resamples: int | None = None,
    seed: int | None = None,
    boost: bool = True,
    lower: float | None = None,
    upper: float | None = None,
    beta: float | None = None,
    precision: float | None = None,
) -> FDPCertificate

Compute p-values once and return a simultaneous FDP certificate.

Supports unweighted empirical Split inference, including detached calibration. Choose the method before inspecting its curve and keep the testing family fixed. Confidence is coverage, not an FDR target. Scientific exchangeability remains the caller's responsibility.

Options match :meth:nonconform.fdr.FDPCertificate.from_p_values. The seed controls certificate Monte Carlo sampling only and does not inherit the fitting seed. The returned certificate is independent of subsequent detector operations; its select() returns a NumPy mask.

Source code in nonconform/detector.py
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
def fdp_bounds(
    self,
    x: pd.DataFrame | pd.Series | np.ndarray,
    *,
    confidence: float = 0.95,
    method: str = "mc_thc",
    n_resamples: int | None = None,
    seed: int | None = None,
    boost: bool = True,
    lower: float | None = None,
    upper: float | None = None,
    beta: float | None = None,
    precision: float | None = None,
) -> FDPCertificate:
    """Compute p-values once and return a simultaneous FDP certificate.

    Supports unweighted empirical Split inference, including detached
    calibration. Choose the method before inspecting its curve and keep
    the testing family fixed. Confidence is coverage, not an FDR target.
    Scientific exchangeability remains the caller's responsibility.

    Options match :meth:`nonconform.fdr.FDPCertificate.from_p_values`.
    The seed controls certificate Monte Carlo sampling only and does not
    inherit the fitting seed. The returned certificate is independent of
    subsequent detector operations; its select() returns a NumPy mask.
    """
    from nonconform._internal.fdp_bounds import validate_scope

    if not self.is_fitted:
        raise NotFittedError("This ConformalDetector instance is not fitted yet.")
    validate_scope(self._result_provenance(None))
    self.compute_p_values(x)
    return self._last_result.fdp_bounds(
        confidence=confidence,
        method=method,
        n_resamples=n_resamples,
        seed=seed,
        boost=boost,
        lower=lower,
        upper=upper,
        beta=beta,
        precision=precision,
    )

Resampling Strategies

nonconform.resampling

Calibration strategies for conformal anomaly detection.

These strategies define how detector replicas and calibration scores are formed. Split uses disjoint fitting and calibration subsets. CrossValidation uses out-of-fold scores, and JackknifeBootstrap uses out-of-bag scores. The latter two are package-specific score-aggregation constructions; their names do not by themselves transfer coverage theorems for conformal prediction intervals to anomaly p-values.

DerandomizedSplits retains separate held-out calibration rows and constructs e-values per split before uniform evidence aggregation and e-BH selection.

Classes:

Name Description
BaseStrategy

Abstract base class for calibration strategies.

Split

Simple train-test split strategy.

DerandomizedSplits

Repeated split-conformal e-values with e-BH selection.

CrossValidation

K-fold cross-validation strategy (includes Jackknife factory).

JackknifeBootstrap

Bootstrap out-of-bag calibration strategy.

BaseStrategy

BaseStrategy(mode: ConformalModeInput = 'plus')

Bases: ABC

Abstract base class for anomaly detection calibration strategies.

This class provides a common interface for various calibration strategies applied to anomaly detectors. Subclasses must implement the core calibration logic and define how calibration data is identified and used.

Attributes:

Name Type Description
_mode ConformalMode

Model retention mode controlling calibration/inference behavior.

Parameters:

Name Type Description Default
mode ConformalModeInput

Model retention mode ("plus" or "single_model"). Equivalent ConformalMode enum values are also accepted.

'plus'
Source code in nonconform/resampling.py
104
105
106
107
108
109
110
111
112
def __init__(self, mode: ConformalModeInput = "plus") -> None:
    """Initialize the base calibration strategy.

    Args:
        mode: Model retention mode (`"plus"` or `"single_model"`).
            Equivalent ``ConformalMode`` enum values are also accepted.
    """
    self._mode: ConformalMode = _normalize_mode(mode)
    self._calibration_ids: list[int] = []
calibration_ids abstractmethod property
calibration_ids: list[int] | None

Indices of data points used for calibration.

fit_calibrate abstractmethod
fit_calibrate(
    x: DataFrame | ndarray,
    detector: AnomalyDetector,
    seed: int | None = None,
    weighted: bool = False,
) -> tuple[list[AnomalyDetector], np.ndarray]

Fits the detector and performs calibration.

Parameters:

Name Type Description Default
x DataFrame | ndarray

The input data for fitting and calibration.

required
detector AnomalyDetector

The anomaly detection model to be fitted and calibrated.

required
seed int | None

Random seed for reproducibility. Defaults to None.

None
weighted bool

Whether to use weighted approach. Defaults to False.

False

Returns:

Type Description
tuple[list[AnomalyDetector], ndarray]

Tuple of (list of trained detectors, calibration scores array).

Source code in nonconform/resampling.py
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
@abc.abstractmethod
def fit_calibrate(
    self,
    x: pd.DataFrame | np.ndarray,
    detector: AnomalyDetector,
    seed: int | None = None,
    weighted: bool = False,
) -> tuple[list[AnomalyDetector], np.ndarray]:
    """Fits the detector and performs calibration.

    Args:
        x: The input data for fitting and calibration.
        detector: The anomaly detection model to be fitted and calibrated.
        seed: Random seed for reproducibility. Defaults to None.
        weighted: Whether to use weighted approach. Defaults to False.

    Returns:
        Tuple of (list of trained detectors, calibration scores array).
    """
    raise NotImplementedError(
        "The fit_calibrate() method must be implemented by subclasses."
    )

Split

Split(n_calib: float | int = 0.1)

Bases: BaseStrategy

Split conformal strategy for fast anomaly detection.

Implements the classical split conformal approach by dividing training data into separate fitting and calibration sets.

Parameters:

Name Type Description Default
n_calib float | int

Size or proportion of data used for calibration. If float, must be between 0.0 and 1.0 (proportion). If int, the absolute number of samples. Defaults to 0.1.

0.1

Examples:

from nonconform import Split

# Use 20% of data for calibration
strategy = Split(n_calib=0.2)

# Use exactly 1000 samples for calibration
strategy = Split(n_calib=1000)
Source code in nonconform/resampling.py
167
168
169
170
def __init__(self, n_calib: float | int = 0.1) -> None:
    super().__init__()
    self._calib_size: float | int = n_calib
    self._calibration_ids: list[int] | None = None
calibration_ids property
calibration_ids: list[int] | None

Indices of calibration samples (None if weighted=False).

calib_size property
calib_size: float | int

Returns the calibration size or proportion.

fit_calibrate
fit_calibrate(
    x: DataFrame | ndarray,
    detector: AnomalyDetector,
    weighted: bool = False,
    seed: int | None = None,
) -> tuple[list[AnomalyDetector], np.ndarray]

Fits detector and generates calibration scores using a data split.

Parameters:

Name Type Description Default
x DataFrame | ndarray

The input data.

required
detector AnomalyDetector

The detector instance to train.

required
weighted bool

If True, stores calibration sample indices. Defaults to False.

False
seed int | None

Random seed for reproducibility. Defaults to None.

None

Returns:

Type Description
tuple[list[AnomalyDetector], ndarray]

Tuple of (list with trained detector, calibration scores array).

Source code in nonconform/resampling.py
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
@ensure_numpy_array
def fit_calibrate(
    self,
    x: pd.DataFrame | np.ndarray,
    detector: AnomalyDetector,
    weighted: bool = False,
    seed: int | None = None,
) -> tuple[list[AnomalyDetector], np.ndarray]:
    """Fits detector and generates calibration scores using a data split.

    Args:
        x: The input data.
        detector: The detector instance to train.
        weighted: If True, stores calibration sample indices. Defaults to False.
        seed: Random seed for reproducibility. Defaults to None.

    Returns:
        Tuple of (list with trained detector, calibration scores array).
    """
    self._validate_n_calib(len(x))
    x_id = np.arange(len(x))
    train_id, calib_id = train_test_split(
        x_id, test_size=self._calib_size, shuffle=True, random_state=seed
    )

    _try_set_random_state(detector, seed)

    detector.fit(x[train_id])
    calibration_set = detector.decision_function(x[calib_id])

    if weighted:
        self._calibration_ids = calib_id.tolist()
    else:
        self._calibration_ids = None
    return [detector], calibration_set

DerandomizedSplits

DerandomizedSplits(
    n_repetitions: int = 5,
    n_calib: float | int = 0.1,
    *,
    alpha_bh: float | None = None,
    tie_seed: int | None = None,
)

Bases: BaseStrategy

Aggregate conformal e-values across repeated random splits with e-BH.

Each replica retains its own held-out calibration scores. Selection converts each replica's scores to e-values before uniformly averaging the evidence; it never pools calibration scores or aggregates raw scores first.

Parameters:

Name Type Description Default
n_repetitions int

Positive number of splits. Defaults to five.

5
n_calib float | int

Calibration count or fraction, with the same meaning as Split.

0.1
alpha_bh float | None

Fixed inner threshold in (0, 1), or None for alpha / 10 using the target supplied to detector.select(). Choose before inspecting the test evidence.

None
tie_seed int | None

Optional non-negative override for randomized score ties. None automatically derives a separate random stream during fit.

None

Examples:

from sklearn.ensemble import IsolationForest
from nonconform import ConformalDetector, DerandomizedSplits

detector = ConformalDetector(
    detector=IsolationForest(),
    strategy=DerandomizedSplits(n_repetitions=5, n_calib=0.2),
    seed=42,
)
# detector.fit(x_reference)
# selected = detector.select(x_test, alpha=0.05)
# evidence = detector.last_selection_result.e_values
Note

Requires unweighted, integrated splits and exchangeable normal reference and null test observations. All repetitions score the same fixed test family. The aggregate null-evidence condition supports one final e-BH application; individual values need not be ordinary e-values. Repetition reduces dependence on a particular split but does not remove randomness.

Source code in nonconform/resampling.py
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
def __init__(
    self,
    n_repetitions: int = 5,
    n_calib: float | int = 0.1,
    *,
    alpha_bh: float | None = None,
    tie_seed: int | None = None,
) -> None:
    super().__init__()
    validate_positive_integer("n_repetitions", n_repetitions)
    if alpha_bh is not None:
        validate_probability("alpha_bh", alpha_bh)
    validate_optional_seed("tie_seed", tie_seed)
    self.n_repetitions = n_repetitions
    self.n_calib = n_calib
    self.alpha_bh = alpha_bh
    self.tie_seed = tie_seed
    self._effective_tie_seed: int | None = None
calibration_ids property
calibration_ids: None

No pooled calibration sample exists for this unweighted strategy.

get_params
get_params(deep: bool = True) -> dict[str, Any]

Return constructor parameters for inspection and sklearn cloning.

Source code in nonconform/resampling.py
309
310
311
312
313
314
315
316
def get_params(self, deep: bool = True) -> dict[str, Any]:
    """Return constructor parameters for inspection and sklearn cloning."""
    return {
        "n_repetitions": self.n_repetitions,
        "n_calib": self.n_calib,
        "alpha_bh": self.alpha_bh,
        "tie_seed": self.tie_seed,
    }
set_params
set_params(**params: Any) -> Self

Update constructor parameters and clear derived random state.

Source code in nonconform/resampling.py
318
319
320
321
322
323
324
325
326
327
328
def set_params(self, **params: Any) -> Self:
    """Update constructor parameters and clear derived random state."""
    current = self.get_params()
    unknown = params.keys() - current.keys()
    if unknown:
        raise ValueError(
            f"Invalid DerandomizedSplits parameters: {sorted(unknown)}"
        )
    updated = type(self)(**(current | params))
    self.__dict__.update(updated.__dict__)
    return self
fit_calibrate
fit_calibrate(
    x: DataFrame | ndarray,
    detector: AnomalyDetector,
    seed: int | None = None,
    weighted: bool = False,
) -> tuple[list[AnomalyDetector], np.ndarray]

Fit independent replicas and return aligned calibration-score rows.

Returns:

Type Description
list[AnomalyDetector]

Models and scores shaped (n_repetitions, n_calibration). Row i

ndarray

contains only held-out scores produced by model i.

Source code in nonconform/resampling.py
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
@ensure_numpy_array
def fit_calibrate(
    self,
    x: pd.DataFrame | np.ndarray,
    detector: AnomalyDetector,
    seed: int | None = None,
    weighted: bool = False,
) -> tuple[list[AnomalyDetector], np.ndarray]:
    """Fit independent replicas and return aligned calibration-score rows.

    Returns:
        Models and scores shaped (n_repetitions, n_calibration). Row i
        contains only held-out scores produced by model i.
    """
    self._effective_tie_seed = None
    x = np.asarray(x)
    if x.ndim != 2 or x.shape[1] == 0:
        raise ValueError(
            "x must be a two-dimensional reference batch with features."
        )
    if weighted:
        raise ValueError(
            "DerandomizedSplits does not support weighted calibration."
        )
    validate_optional_seed("seed", seed)
    validate_positive_integer("n_repetitions", self.n_repetitions)
    split = Split(n_calib=self.n_calib)
    split._validate_n_calib(len(x))
    n_calibration = (
        math.ceil(len(x) * self.n_calib)
        if isinstance(self.n_calib, float)
        else self.n_calib
    )
    split_stream, tie_stream = np.random.SeedSequence(seed).spawn(2)
    models = []
    calibration_rows = []
    for stream in split_stream.spawn(self.n_repetitions):
        split_seed = int(stream.generate_state(1)[0])
        replica = set_params(deepcopy(detector), split_seed)
        fitted, scores = split.fit_calibrate(x, replica, seed=split_seed)
        row = as_1d_numeric("calibration scores", scores)
        if row.size != n_calibration:
            raise ValueError(
                "Each model must return one score per calibration row."
            )
        validate_finite("calibration scores", row)
        models.append(fitted[0])
        calibration_rows.append(row.copy())
    calibration_scores = np.vstack(calibration_rows)
    self._effective_tie_seed = (
        int(tie_stream.generate_state(1, dtype=np.uint64)[0])
        if self.tie_seed is None
        else self.tie_seed
    )
    return models, calibration_scores

CrossValidation

CrossValidation(
    k: int | None = 5,
    mode: ConformalModeInput = "plus",
    shuffle: bool = True,
)

Bases: BaseStrategy

K-fold out-of-fold calibration for conformal anomaly scoring.

The strategy trains one detector per fold and records scores for observations while they are held out. In "plus" mode, test scores are aggregated over the retained fold models before comparison with the out-of-fold calibration scores. This is the package's anomaly-score construction, not a claim that a CV+ prediction-interval theorem applies unchanged.

Parameters:

Name Type Description Default
k int | None

Number of folds. If None, uses leave-one-out (k=n at fit time).

5
mode ConformalModeInput

Model retention mode ("plus" or "single_model"). Equivalent ConformalMode values are accepted. Defaults to "plus".

'plus'
shuffle bool

Whether to shuffle data before splitting. Defaults to True. Set to False for deterministic leave-one-out (Jackknife).

True

Examples:

from nonconform import CrossValidation

# 5-fold cross-validation
strategy = CrossValidation(k=5)

# Leave-one-out (Jackknife) via factory
strategy = CrossValidation.jackknife()
Source code in nonconform/resampling.py
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
def __init__(
    self,
    k: int | None = 5,
    mode: ConformalModeInput = "plus",
    shuffle: bool = True,
) -> None:
    super().__init__(mode)
    if not isinstance(shuffle, bool):
        raise TypeError(
            f"shuffle must be a boolean value, got {type(shuffle).__name__}."
        )
    self._k: int | None = k
    self._shuffle: bool = shuffle
    self._is_jackknife = k is None

    # Warn if using single-model mode
    if self._mode is ConformalMode.SINGLE_MODEL:
        _crossval_logger.warning(
            "Setting mode=ConformalMode.SINGLE_MODEL may compromise conformal "
            "validity. mode=ConformalMode.PLUS is recommended."
        )

    self._detector_list: list[AnomalyDetector] = []
    self._calibration_set: np.ndarray = np.array([])
    self._calibration_ids: list[int] = []
calibration_ids property
calibration_ids: list[int]

Indices of samples used for calibration.

k property
k: int | None

Number of folds (None for jackknife mode).

mode property
mode: Literal['plus', 'single_model']

User-facing model retention mode.

jackknife classmethod
jackknife(
    mode: ConformalModeInput = "plus",
) -> CrossValidation

Create Leave-One-Out cross-validation (deterministic, no shuffle).

This factory method creates a Jackknife strategy, which is a special case of k-fold CV where k equals n (the dataset size). Each sample is left out exactly once for calibration.

Parameters:

Name Type Description Default
mode ConformalModeInput

Model retention mode ("plus" or "single_model").

'plus'

Returns:

Type Description
CrossValidation

CrossValidation configured for leave-one-out.

Examples:

from nonconform import CrossValidation

strategy = CrossValidation.jackknife()
print(strategy.k, strategy.mode)
Source code in nonconform/resampling.py
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
@classmethod
def jackknife(cls, mode: ConformalModeInput = "plus") -> CrossValidation:
    """Create Leave-One-Out cross-validation (deterministic, no shuffle).

    This factory method creates a Jackknife strategy, which is a special
    case of k-fold CV where k equals n (the dataset size). Each sample is
    left out exactly once for calibration.

    Args:
        mode: Model retention mode (`"plus"` or `"single_model"`).

    Returns:
        CrossValidation configured for leave-one-out.

    Examples:
        ```python
        from nonconform import CrossValidation

        strategy = CrossValidation.jackknife()
        print(strategy.k, strategy.mode)
        ```
    """
    return cls(k=None, mode=mode, shuffle=False)
fit_calibrate
fit_calibrate(
    x: DataFrame | ndarray,
    detector: AnomalyDetector,
    seed: int | None = None,
    weighted: bool = False,
) -> tuple[list[AnomalyDetector], np.ndarray]

Fit and calibrate using k-fold cross-validation.

Parameters:

Name Type Description Default
x DataFrame | ndarray

Input data matrix.

required
detector AnomalyDetector

The base anomaly detector.

required
seed int | None

Random seed for reproducibility. Defaults to None.

None
weighted bool

Whether to use weighted calibration. Defaults to False.

False

Returns:

Type Description
tuple[list[AnomalyDetector], ndarray]

Tuple of (list of trained detectors, calibration scores array).

Raises:

Type Description
ValueError

If k < 2 or not enough samples for specified k.

Source code in nonconform/resampling.py
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
@ensure_numpy_array
def fit_calibrate(
    self,
    x: pd.DataFrame | np.ndarray,
    detector: AnomalyDetector,
    seed: int | None = None,
    weighted: bool = False,
) -> tuple[list[AnomalyDetector], np.ndarray]:
    """Fit and calibrate using k-fold cross-validation.

    Args:
        x: Input data matrix.
        detector: The base anomaly detector.
        seed: Random seed for reproducibility. Defaults to None.
        weighted: Whether to use weighted calibration. Defaults to False.

    Returns:
        Tuple of (list of trained detectors, calibration scores array).

    Raises:
        ValueError: If k < 2 or not enough samples for specified k.
    """
    self._detector_list.clear()
    self._calibration_ids = []

    detector_ = detector
    n_samples = len(x)

    # Determine k (for jackknife mode, k=n)
    k = n_samples if self._is_jackknife else self._k

    if k < 2:
        exc = ValueError(
            f"k must be at least 2 for k-fold cross-validation, got {k}"
        )
        exc.add_note(f"Received k={k}, which is invalid.")
        exc.add_note(
            "Cross-validation requires at least one split for training "
            "and one for calibration."
        )
        raise exc

    if n_samples < k:
        exc = ValueError(
            f"Not enough samples ({n_samples}) for "
            f"k-fold cross-validation with k={k}"
        )
        exc.add_note(f"Each fold needs at least 1 sample, but {n_samples} < {k}.")
        raise exc

    self._calibration_set = np.empty(n_samples, dtype=np.float64)
    calibration_offset = 0

    folds = KFold(
        n_splits=k,
        shuffle=self._shuffle,
        random_state=seed if self._shuffle else None,
    )

    fold_iterator = (
        tqdm(folds.split(x), total=k, desc="Calibration")
        if _crossval_logger.isEnabledFor(logging.INFO)
        else folds.split(x)
    )

    for i, (train_idx, calib_idx) in enumerate(fold_iterator):
        self._calibration_ids.extend(calib_idx.tolist())

        model = copy(detector_)
        _try_set_random_state(model, seed)
        model.fit(x[train_idx])

        if self._mode is ConformalMode.PLUS:
            self._detector_list.append(deepcopy(model))

        fold_scores = model.decision_function(x[calib_idx])
        n_fold_samples = len(fold_scores)
        end_idx = calibration_offset + n_fold_samples
        self._calibration_set[calibration_offset:end_idx] = fold_scores
        calibration_offset += n_fold_samples

    if self._mode is ConformalMode.SINGLE_MODEL:
        model = copy(detector_)
        _try_set_random_state(model, seed)
        model.fit(x)
        self._detector_list.append(deepcopy(model))

    return self._detector_list, self._calibration_set

JackknifeBootstrap

JackknifeBootstrap(
    n_bootstraps: int = 100,
    aggregation_method: BootstrapAggregationMethod = "mean",
    mode: ConformalModeInput = "plus",
)

Bases: BaseStrategy

Bootstrap and out-of-bag calibration for conformal anomaly scoring.

Each bootstrap replica is fitted on a sample drawn with replacement. Every reference observation receives a calibration score aggregated over replicas for which that observation was out of bag. In "plus" mode, test scores are aggregated over all retained replicas.

The construction is inspired by jackknife+-after-bootstrap (JaB+), but this class produces anomaly scores and p-values rather than the predictive intervals studied by the JaB+ theorem. Do not infer an interval-coverage guarantee solely from the class name.

Parameters:

Name Type Description Default
n_bootstraps int

Number of bootstrap iterations. Defaults to 100.

100
aggregation_method BootstrapAggregationMethod

How to aggregate OOB predictions ("mean" or "median"). Defaults to "mean".

'mean'
mode ConformalModeInput

Model retention mode ("plus" or "single_model"). Equivalent ConformalMode values are accepted. Defaults to "plus".

'plus'
References

Kim, Byol, Chen Xu, and Rina Foygel Barber. "Predictive Inference Is Free with the Jackknife+-after-Bootstrap." NeurIPS 2020.

Source code in nonconform/resampling.py
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
def __init__(
    self,
    n_bootstraps: int = 100,
    aggregation_method: BootstrapAggregationMethod = "mean",
    mode: ConformalModeInput = "plus",
) -> None:
    super().__init__(mode=mode)

    if n_bootstraps < 2:
        exc = ValueError(
            f"Number of bootstraps must be at least 2, got {n_bootstraps}."
        )
        exc.add_note(f"Received n_bootstraps={n_bootstraps}, which is invalid.")
        raise exc

    normalized_aggregation_method = normalize_bootstrap_aggregation_method(
        aggregation_method
    )

    if self._mode is ConformalMode.SINGLE_MODEL:
        _bootstrap_logger.warning(
            "Setting mode=ConformalMode.SINGLE_MODEL may compromise conformal "
            "validity. mode=ConformalMode.PLUS is recommended."
        )

    self._n_bootstraps: int = n_bootstraps
    self._aggregation_method: BootstrapAggregationMethod = (
        normalized_aggregation_method
    )

    self._detector_list: list[AnomalyDetector] = []
    self._calibration_set: np.ndarray = np.array([])
    self._calibration_ids: list[int] = []

    # Internal state
    self._bootstrap_models: list[AnomalyDetector | None] = []
    self._oob_mask: np.ndarray = np.array([])
calibration_ids property
calibration_ids: list[int]

Indices used for calibration (all samples in JaB+).

n_bootstraps property
n_bootstraps: int

Number of bootstrap iterations.

aggregation_method property
aggregation_method: BootstrapAggregationMethod

Aggregation method for OOB predictions.

fit_calibrate
fit_calibrate(
    x: DataFrame | ndarray,
    detector: AnomalyDetector,
    seed: int | None = None,
    weighted: bool = False,
    n_jobs: int | None = None,
) -> tuple[list[AnomalyDetector], np.ndarray]

Fit bootstrap replicas and compute out-of-bag calibration scores.

Parameters:

Name Type Description Default
x DataFrame | ndarray

Input data matrix.

required
detector AnomalyDetector

The base anomaly detector.

required
seed int | None

Random seed for reproducibility. Defaults to None.

None
weighted bool

Accepted for the shared strategy interface. Calibration indices already include every input row in this strategy.

False
n_jobs int | None

Number of parallel jobs. Use -1 for all available cores. Defaults to None (sequential).

None

Returns:

Type Description
tuple[list[AnomalyDetector], ndarray]

Tuple of (list of trained detectors, calibration scores array).

Source code in nonconform/resampling.py
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
@ensure_numpy_array
def fit_calibrate(
    self,
    x: pd.DataFrame | np.ndarray,
    detector: AnomalyDetector,
    seed: int | None = None,
    weighted: bool = False,
    n_jobs: int | None = None,
) -> tuple[list[AnomalyDetector], np.ndarray]:
    """Fit bootstrap replicas and compute out-of-bag calibration scores.

    Args:
        x: Input data matrix.
        detector: The base anomaly detector.
        seed: Random seed for reproducibility. Defaults to None.
        weighted: Accepted for the shared strategy interface. Calibration
            indices already include every input row in this strategy.
        n_jobs: Number of parallel jobs. Use -1 for all available cores.
            Defaults to None (sequential).

    Returns:
        Tuple of (list of trained detectors, calibration scores array).
    """
    n_samples = len(x)
    generator = np.random.default_rng(seed)

    _bootstrap_logger.info(
        f"Bootstrap (JaB+): {n_samples:,} samples, "
        f"{self._n_bootstraps:,} iterations"
    )

    self._bootstrap_models = [None] * self._n_bootstraps
    all_bootstrap_indices, self._oob_mask = self._generate_bootstrap_indices(
        generator, n_samples
    )

    if n_jobs == -1:
        n_jobs = os.cpu_count() or 1
    elif n_jobs is not None and n_jobs < 1:
        raise ValueError(
            f"n_jobs must be None, -1, or a positive integer; got {n_jobs}."
        )

    if n_jobs is None or n_jobs == 1:
        bootstrap_iterator = (
            tqdm(range(self._n_bootstraps), desc="Calibration")
            if _bootstrap_logger.isEnabledFor(logging.INFO)
            else range(self._n_bootstraps)
        )
        for i in bootstrap_iterator:
            bootstrap_indices = all_bootstrap_indices[i]
            model = _train_bootstrap_model(detector, x, bootstrap_indices, seed)
            self._bootstrap_models[i] = model
    else:
        self._train_models_parallel(
            detector, x, all_bootstrap_indices, seed, n_jobs
        )

    oob_scores = self._compute_oob_scores(x)

    self._calibration_set = oob_scores
    self._calibration_ids = list(range(n_samples))

    if self._mode is ConformalMode.PLUS:
        self._detector_list = self._bootstrap_models.copy()
    else:
        final_model = deepcopy(detector)
        _try_set_random_state(final_model, seed)
        final_model.fit(x)
        self._detector_list = [final_model]

    return self._detector_list, self._calibration_set

P-Value Estimation

nonconform.scoring

Tail-probability estimation strategies for calibrated anomaly scores.

Empirical implements rank-based conformal p-values. ConditionalEmpirical adds a calibration map for a stronger conditional target. Probabilistic instead estimates a smooth score-tail probability with a kernel density model; it does not inherit the exact finite-sample guarantee of the empirical rank.

Classes:

Name Description
BaseEstimation

Abstract base class for p-value estimation.

Empirical

Classical empirical p-value estimation using discrete CDF.

ConditionalEmpirical

Conditionally calibrated empirical p-values.

Probabilistic

KDE-based score-tail probability estimation.

Kernel

Bases: Enum

Kernel functions for KDE-based score-tail estimation.

Attributes:

Name Type Description
GAUSSIAN

Gaussian (normal) kernel.

EXPONENTIAL

Exponential kernel.

BOX

Box (uniform) kernel.

TRIANGULAR

Triangular kernel.

EPANECHNIKOV

Epanechnikov kernel.

BIWEIGHT

Biweight (quartic) kernel.

TRIWEIGHT

Triweight kernel.

TRICUBE

Tricube kernel.

COSINE

Cosine kernel.

BaseEstimation

Bases: ABC

Abstract base for p-value and score-tail estimation strategies.

compute_p_values abstractmethod
compute_p_values(
    scores: ndarray,
    calibration_set: ndarray,
    weights: tuple[ndarray, ndarray] | None = None,
) -> np.ndarray

Compute p-values or strategy-defined tail estimates for test scores.

Parameters:

Name Type Description Default
scores ndarray

Test instance anomaly scores (1D array).

required
calibration_set ndarray

Calibration anomaly scores (1D array).

required
weights tuple[ndarray, ndarray] | None

Optional (w_calib, w_test) tuple for weighted conformal.

None

Returns:

Type Description
ndarray

Array of values for each test instance. The concrete strategy

ndarray

defines their statistical interpretation.

Source code in nonconform/scoring.py
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
@abstractmethod
def compute_p_values(
    self,
    scores: np.ndarray,
    calibration_set: np.ndarray,
    weights: tuple[np.ndarray, np.ndarray] | None = None,
) -> np.ndarray:
    """Compute p-values or strategy-defined tail estimates for test scores.

    Args:
        scores: Test instance anomaly scores (1D array).
        calibration_set: Calibration anomaly scores (1D array).
        weights: Optional (w_calib, w_test) tuple for weighted conformal.

    Returns:
        Array of values for each test instance. The concrete strategy
        defines their statistical interpretation.
    """
    pass
get_metadata
get_metadata() -> dict[str, Any]

Optional auxiliary data exposed after compute_p_values.

Source code in nonconform/scoring.py
86
87
88
def get_metadata(self) -> dict[str, Any]:
    """Optional auxiliary data exposed after compute_p_values."""
    return {}
set_seed
set_seed(seed: int | None) -> None

Set random seed for reproducibility.

Parameters:

Name Type Description Default
seed int | None

Random seed value or None.

required
Source code in nonconform/scoring.py
90
91
92
93
94
95
96
97
def set_seed(self, seed: int | None) -> None:
    """Set random seed for reproducibility.

    Args:
        seed: Random seed value or None.
    """
    if hasattr(self, "_seed"):
        self._seed = seed

Empirical

Empirical(tie_break: TieBreakModeInput = 'classical')

Bases: BaseEstimation

Classical rank-based empirical conformal p-values.

Computes p-values using deterministic tie handling by default. Optionally supports randomized smoothing of tied calibration mass and the test point's own mass, eliminating the classical resolution floor (Jin & Candes 2023).

Parameters:

Name Type Description Default
tie_break TieBreakModeInput

Tie-breaking strategy ("classical" or "randomized"). Equivalent TieBreakMode enum values are also accepted.

'classical'

Examples:

import numpy as np

from nonconform import Empirical

calibration_scores = np.array([0.2, 0.5, 0.7, 1.1, 1.4])
test_scores = np.array([0.6, 1.5])
estimation = Empirical()  # tie_break="classical" by default
p_values = estimation.compute_p_values(test_scores, calibration_scores)
print(p_values)

# For randomized smoothing:
randomized = Empirical(tie_break="randomized")
randomized.set_seed(42)
print(randomized.compute_p_values(test_scores, calibration_scores))
Source code in nonconform/scoring.py
130
131
132
def __init__(self, tie_break: TieBreakModeInput = "classical") -> None:
    self._tie_break = _normalize_tie_break_mode(tie_break)
    self._seed: int | None = None
set_seed
set_seed(seed: int | None) -> None

Set random seed for reproducibility.

Source code in nonconform/scoring.py
134
135
136
def set_seed(self, seed: int | None) -> None:
    """Set random seed for reproducibility."""
    self._seed = seed
compute_p_values
compute_p_values(
    scores: ndarray,
    calibration_set: ndarray,
    weights: tuple[ndarray, ndarray] | None = None,
) -> np.ndarray

Compute empirical p-values from calibration set.

Source code in nonconform/scoring.py
138
139
140
141
142
143
144
145
146
147
148
149
def compute_p_values(
    self,
    scores: np.ndarray,
    calibration_set: np.ndarray,
    weights: tuple[np.ndarray, np.ndarray] | None = None,
) -> np.ndarray:
    """Compute empirical p-values from calibration set."""
    randomized = self._tie_break is TieBreakMode.RANDOMIZED
    rng = np.random.default_rng(self._seed) if randomized else None
    if weights is not None:
        return self._compute_weighted(scores, calibration_set, weights, rng)
    return self._compute_standard(scores, calibration_set, rng)

ConditionalEmpirical

ConditionalEmpirical(
    *,
    delta: float = 0.05,
    method: str | ConditionalCalibrationMethod = "mc",
    tie_break: TieBreakModeInput = "classical",
    simes_kden: int = 2,
    mc_num_simulations: int = 10000,
)

Bases: Empirical

Conditionally calibrated empirical conformal p-values (CCCPV).

This estimator first computes classical empirical conformal p-values and then applies a finite-sample calibration map:

.. math:: p_j = \frac{1 + \sum_{i=1}^{n_{\text{cal}}}\mathbf{1}[s_i \ge s_j]} {n_{\text{cal}} + 1}, \qquad \tilde p_j = C_{n_{\text{cal}},\delta}(p_j).

Supported calibration maps are "mc", "simes", "dkwm", and "asymptotic".

References

Bates et al. (2023), Testing for outliers with conformal p-values. Reference implementation: https://github.com/msesia/conditional-conformal-pvalues

Note

Weighted conformal p-values are intentionally not supported in this first release of ConditionalEmpirical.

Parameters:

Name Type Description Default
delta float

Failure-probability parameter used by the conditional calibration map; the associated calibration event has probability at least 1 - delta under the method's assumptions. Must be in (0, 1). Defaults to 0.05.

0.05
method str | ConditionalCalibrationMethod

Conditional calibration method. One of {"mc", "simes", "dkwm", "asymptotic"}. Defaults to "mc".

'mc'
tie_break TieBreakModeInput

Tie-breaking strategy used for base empirical p-values ("classical" or "randomized").

'classical'
simes_kden int

Denominator used to derive k = floor(n_cal / simes_kden) for the Simes calibration map. Must be a positive integer. Defaults to 2.

2
mc_num_simulations int

Monte Carlo sample size used to estimate the finite-sample correction for method="mc". Defaults to 10,000.

10000
Source code in nonconform/scoring.py
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
def __init__(
    self,
    *,
    delta: float = 0.05,
    method: str | ConditionalCalibrationMethod = "mc",
    tie_break: TieBreakModeInput = "classical",
    simes_kden: int = 2,
    mc_num_simulations: int = 10_000,
) -> None:
    super().__init__(tie_break=tie_break)
    try:
        delta_float = float(delta)
    except (TypeError, ValueError) as exc:
        raise ValueError("delta must be a float in (0, 1).") from exc
    if not np.isfinite(delta_float) or not (0.0 < delta_float < 1.0):
        raise ValueError(f"delta must be in (0, 1), got {delta!r}.")
    if (
        isinstance(simes_kden, bool)
        or not isinstance(simes_kden, int)
        or simes_kden < 1
    ):
        raise ValueError("simes_kden must be a positive integer.")
    if (
        isinstance(mc_num_simulations, bool)
        or not isinstance(mc_num_simulations, int)
        or mc_num_simulations < 100
    ):
        raise ValueError("mc_num_simulations must be an integer >= 100.")

    self._delta = delta_float
    self._method = normalize_conditional_calibration_method(method)
    self._simes_kden = simes_kden
    self._mc_num_simulations = mc_num_simulations
    self._mc_correction_cache: dict[tuple[int, float], float] = {}
set_seed
set_seed(seed: int | None) -> None

Set random seed for reproducibility.

Source code in nonconform/scoring.py
256
257
258
259
260
def set_seed(self, seed: int | None) -> None:
    """Set random seed for reproducibility."""
    super().set_seed(seed)
    # MC correction estimation depends on RNG; invalidate cached estimates.
    self._mc_correction_cache.clear()
compute_p_values
compute_p_values(
    scores: ndarray,
    calibration_set: ndarray,
    weights: tuple[ndarray, ndarray] | None = None,
) -> np.ndarray

Compute conditionally calibrated conformal p-values.

Source code in nonconform/scoring.py
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
def compute_p_values(
    self,
    scores: np.ndarray,
    calibration_set: np.ndarray,
    weights: tuple[np.ndarray, np.ndarray] | None = None,
) -> np.ndarray:
    """Compute conditionally calibrated conformal p-values."""
    if weights is not None:
        raise ValueError(
            "ConditionalEmpirical does not support weighted p-values. "
            "Use Empirical or Probabilistic for weighted conformal mode."
        )

    base_p = super().compute_p_values(scores, calibration_set, weights=None)
    n_cal = len(np.asarray(calibration_set).ravel())

    cache_key = (n_cal, self._delta)
    cached_fs = (
        self._mc_correction_cache.get(cache_key) if self._method == "mc" else None
    )
    rng = np.random.default_rng(self._seed) if self._seed is not None else None
    calibrated, fs_correction = calibrate_conditional_p_values(
        base_p,
        n_calibration=n_cal,
        delta=self._delta,
        method=self._method,
        simes_kden=self._simes_kden,
        fs_correction=cached_fs,
        rng=rng,
        mc_num_simulations=self._mc_num_simulations,
    )
    if self._method == "mc" and fs_correction is not None:
        self._mc_correction_cache[cache_key] = fs_correction
    return calibrated

Probabilistic

Probabilistic(
    kernel: Kernel | Sequence[Kernel] = Kernel.GAUSSIAN,
    n_trials: int = 100,
    cv_folds: int = -1,
)

Bases: BaseEstimation

KDE-based continuous estimates of score-tail probability.

Fits a kernel density estimate to calibration scores and evaluates its survival function. The results are model-based estimates, not rank-based conformal p-values, and therefore do not have the empirical estimator's exact finite-sample distribution-free guarantee. Smoothness or finer numerical resolution is not evidence of calibration.

The estimator supports automatic hyperparameter tuning and calibration weights. In weighted mode, only calibration weights are applied to the KDE; test weights are intentionally not injected into the survival calculation. This is not the exact discrete weighted conformal construction, and it lets tail estimates reach 0 instead of imposing the lower bound w_test / (sum_calib_weight + w_test) that the discrete weighted formula would impose.

Parameters:

Name Type Description Default
kernel Kernel | Sequence[Kernel]

Kernel function or list (list triggers kernel tuning). Bandwidth is always auto-tuned. Defaults to Kernel.GAUSSIAN.

GAUSSIAN
n_trials int

Number of Optuna trials for tuning. Defaults to 100.

100
cv_folds int

CV folds for tuning (-1 for leave-one-out). Defaults to -1.

-1

Examples:

import numpy as np

from nonconform import Probabilistic
from nonconform.enums import Kernel

rng = np.random.default_rng(42)
calibration_scores = rng.normal(size=200)
test_scores = np.array([0.0, 1.0, 2.0])

# n_trials=0 skips optional hyperparameter search for a quick example.
estimation = Probabilistic(kernel=Kernel.GAUSSIAN, n_trials=0)
tail_estimates = estimation.compute_p_values(test_scores, calibration_scores)
print(tail_estimates)
Source code in nonconform/scoring.py
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
def __init__(
    self,
    kernel: Kernel | Sequence[Kernel] = Kernel.GAUSSIAN,
    n_trials: int = 100,
    cv_folds: int = -1,
) -> None:
    self._kernel = kernel
    self._n_trials = n_trials
    self._cv_folds = cv_folds
    self._seed = None

    self._tuned_params: dict | None = None
    self._kde_model = None
    self._calibration_hash: int | None = None
    self._kde_eval_grid: np.ndarray | None = None
    self._kde_cdf_values: np.ndarray | None = None
    self._kde_total_weight: float | None = None
compute_p_values
compute_p_values(
    scores: ndarray,
    calibration_set: ndarray,
    weights: tuple[ndarray, ndarray] | None = None,
) -> np.ndarray

Compute continuous p-values using KDE.

Lazy fitting: tunes and fits KDE on first call or when calibration changes. Note: When weights are provided, this estimator uses only calibration weights to shape the KDE. Test weights are accepted for API parity but do not set a positive lower bound on p-values.

Source code in nonconform/scoring.py
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
def compute_p_values(
    self,
    scores: np.ndarray,
    calibration_set: np.ndarray,
    weights: tuple[np.ndarray, np.ndarray] | None = None,
) -> np.ndarray:
    """Compute continuous p-values using KDE.

    Lazy fitting: tunes and fits KDE on first call or when calibration changes.
    Note: When weights are provided, this estimator uses only calibration
    weights to shape the KDE. Test weights are accepted for API parity but
    do not set a positive lower bound on p-values.
    """
    if weights is not None:
        w_calib, _w_test = weights
    else:
        w_calib, _w_test = None, None

    if weights is None:
        current_hash = hash(calibration_set.tobytes())
    else:
        current_hash = hash((calibration_set.tobytes(), w_calib.tobytes()))

    if self._kde_model is None or self._calibration_hash != current_hash:
        self._fit_kde(calibration_set, w_calib)
        self._calibration_hash = current_hash

    sum_calib_weight = (
        float(np.sum(w_calib))
        if w_calib is not None
        else float(len(calibration_set))
    )

    return self._compute_p_values_from_kde(scores, sum_calib_weight)
get_metadata
get_metadata() -> dict[str, Any]

Return KDE metadata after p-value computation.

Source code in nonconform/scoring.py
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
def get_metadata(self) -> dict[str, Any]:
    """Return KDE metadata after p-value computation."""
    if (
        self._kde_eval_grid is None
        or self._kde_cdf_values is None
        or self._kde_total_weight is None
    ):
        return {}
    return {
        "kde": {
            "eval_grid": self._kde_eval_grid.copy(),
            "cdf_values": self._kde_cdf_values.copy(),
            "total_weight": float(self._kde_total_weight),
        }
    }

calculate_p_val

calculate_p_val(
    scores: ndarray,
    calibration_set: ndarray,
    tie_break: TieBreakModeInput = "classical",
    rng: Generator | None = None,
) -> np.ndarray

Calculate empirical p-values (standalone function).

Uses classical deterministic tie handling by default. Randomized mode interpolates tied calibration mass and the test point's own mass, removing the classical positive resolution floor.

Parameters:

Name Type Description Default
scores ndarray

Test instance anomaly scores (1D array).

required
calibration_set ndarray

Calibration anomaly scores (1D array).

required
tie_break TieBreakModeInput

Tie-breaking strategy for equal scores ("classical" or "randomized"). Equivalent TieBreakMode values are accepted.

'classical'
rng Generator | None

Optional random number generator for reproducibility.

None

Returns:

Type Description
ndarray

Array of p-values for each test instance.

Source code in nonconform/scoring.py
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
def calculate_p_val(
    scores: np.ndarray,
    calibration_set: np.ndarray,
    tie_break: TieBreakModeInput = "classical",
    rng: np.random.Generator | None = None,
) -> np.ndarray:
    """Calculate empirical p-values (standalone function).

    Uses classical deterministic tie handling by default. Randomized mode
    interpolates tied calibration mass and the test point's own mass, removing
    the classical positive resolution floor.

    Args:
        scores: Test instance anomaly scores (1D array).
        calibration_set: Calibration anomaly scores (1D array).
        tie_break: Tie-breaking strategy for equal scores (`"classical"` or
            `"randomized"`). Equivalent `TieBreakMode` values are accepted.
        rng: Optional random number generator for reproducibility.

    Returns:
        Array of p-values for each test instance.
    """
    mode = _normalize_tie_break_mode(tie_break)

    sorted_cal = np.sort(calibration_set)
    n_cal = len(calibration_set)

    if mode is TieBreakMode.CLASSICAL:
        # Deterministic discrete formula: count calibration scores >= score.
        ranks = n_cal - np.searchsorted(sorted_cal, scores, side="left")
        return (1.0 + ranks) / (1.0 + n_cal)

    # Randomized tie handling: separate strictly greater and ties
    pos_right = np.searchsorted(sorted_cal, scores, side="right")
    pos_left = np.searchsorted(sorted_cal, scores, side="left")
    n_greater = n_cal - pos_right  # strictly greater
    n_equal = pos_right - pos_left  # ties

    if rng is None:
        rng = np.random.default_rng()
    u = rng.uniform(size=len(scores))

    return (n_greater + (n_equal + 1) * u) / (1.0 + n_cal)

calculate_weighted_p_val

calculate_weighted_p_val(
    scores: ndarray,
    calibration_set: ndarray,
    test_weights: ndarray,
    calib_weights: ndarray,
    tie_break: TieBreakModeInput = "classical",
    rng: Generator | None = None,
) -> np.ndarray

Calculate weighted empirical p-values (standalone function).

The default "classical" mode uses deterministic discrete tie handling: calibration mass at or above the test score is included, together with the test point's own weight. For calibration mass W_cal and test score s with test weight w_test, this is (W_{>=}(s) + w_test) / (W_cal + w_test). This variant is conservative for discrete scores and agrees with the unweighted classical formula when all weights are one. Without calibration scores tied with s, it equals the strict deterministic formula used by Jin and Candes (2023).

The "randomized" mode uses the tie-safe weighted conformal formula (W_{>}(s) + U * (W_{=}(s) + w_test)) / (W_cal + w_test), where U is uniform on [0, 1]. It avoids the discrete resolution floor and interpolates both tied calibration mass and the test point's own mass.

Parameters:

Name Type Description Default
scores ndarray

Test instance anomaly scores (1D array).

required
calibration_set ndarray

Calibration anomaly scores (1D array).

required
test_weights ndarray

Test instance weights (1D array).

required
calib_weights ndarray

Calibration weights (1D array).

required
tie_break TieBreakModeInput

Tie-breaking strategy for equal scores ("classical" or "randomized"). Equivalent TieBreakMode values are accepted.

'classical'
rng Generator | None

Optional random number generator for reproducibility.

None

Returns:

Type Description
ndarray

Array of weighted p-values for each test instance.

Note

Classical mode has the positive lower bound test_weights / (sum(calib_weights) + test_weights) when no calibration mass lies above the test score. Randomized mode has no such positive floor because it multiplies the test point's mass by U.

Source code in nonconform/scoring.py
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
def calculate_weighted_p_val(
    scores: np.ndarray,
    calibration_set: np.ndarray,
    test_weights: np.ndarray,
    calib_weights: np.ndarray,
    tie_break: TieBreakModeInput = "classical",
    rng: np.random.Generator | None = None,
) -> np.ndarray:
    """Calculate weighted empirical p-values (standalone function).

    The default ``"classical"`` mode uses deterministic discrete tie handling:
    calibration mass at or above the test score is included, together with the
    test point's own weight. For calibration mass ``W_cal`` and test score
    ``s`` with test weight ``w_test``, this is
    ``(W_{>=}(s) + w_test) / (W_cal + w_test)``. This variant is conservative
    for discrete scores and agrees with the unweighted classical formula when
    all weights are one. Without calibration scores tied with ``s``, it equals
    the strict deterministic formula used by Jin and Candes (2023).

    The ``"randomized"`` mode uses the tie-safe weighted conformal formula
    ``(W_{>}(s) + U * (W_{=}(s) + w_test)) / (W_cal + w_test)``, where
    ``U`` is uniform on ``[0, 1]``. It avoids the discrete resolution floor and
    interpolates both tied calibration mass and the test point's own mass.

    Args:
        scores: Test instance anomaly scores (1D array).
        calibration_set: Calibration anomaly scores (1D array).
        test_weights: Test instance weights (1D array).
        calib_weights: Calibration weights (1D array).
        tie_break: Tie-breaking strategy for equal scores (`"classical"` or
            `"randomized"`). Equivalent `TieBreakMode` values are accepted.
        rng: Optional random number generator for reproducibility.

    Returns:
        Array of weighted p-values for each test instance.

    Note:
        Classical mode has the positive lower bound
        ``test_weights / (sum(calib_weights) + test_weights)`` when no
        calibration mass lies above the test score. Randomized mode has no such
        positive floor because it multiplies the test point's mass by ``U``.
    """
    mode = _normalize_tie_break_mode(tie_break)

    try:
        scores = _as_1d("scores", scores).astype(float, copy=False)
        calibration_set = _as_1d("calibration_set", calibration_set).astype(
            float, copy=False
        )
        w_scores = _as_1d("test_weights", test_weights).astype(float, copy=False)
        w_calib = _as_1d("calib_weights", calib_weights).astype(float, copy=False)
    except (TypeError, ValueError) as exc:
        raise ValueError(
            "scores, calibration_set, test_weights, and calib_weights must be numeric."
        ) from exc

    if len(scores) != len(w_scores):
        raise ValueError(
            "scores and test_weights must have the same length. "
            f"Got {len(scores)} and {len(w_scores)}."
        )
    if len(calibration_set) != len(w_calib):
        raise ValueError(
            "calibration_set and calib_weights must have the same length. "
            f"Got {len(calibration_set)} and {len(w_calib)}."
        )
    _validate_finite("scores", scores)
    _validate_finite("calibration_set", calibration_set)
    _validate_finite("test_weights", w_scores)
    _validate_finite("calib_weights", w_calib)
    if np.any(w_scores < 0):
        raise ValueError("test_weights must be non-negative.")
    if np.any(w_calib < 0):
        raise ValueError("calib_weights must be non-negative.")

    sort_idx = np.argsort(calibration_set)
    sorted_scores = calibration_set[sort_idx]
    sorted_weights = w_calib[sort_idx]

    cumulative_weights = np.concatenate(([0.0], np.cumsum(sorted_weights)))
    total_weight = float(cumulative_weights[-1])
    if total_weight <= 0:
        raise ValueError("calib_weights must sum to a positive value.")

    left_idx = np.searchsorted(sorted_scores, scores, side="left")
    right_idx = np.searchsorted(sorted_scores, scores, side="right")

    if mode is TieBreakMode.CLASSICAL:
        weighted_at_or_above = total_weight - cumulative_weights[left_idx]
        numerator = weighted_at_or_above + w_scores
    else:
        weighted_strictly_above = total_weight - cumulative_weights[right_idx]
        weighted_equal = cumulative_weights[right_idx] - cumulative_weights[left_idx]

        if rng is None:
            rng = np.random.default_rng()
        u = rng.uniform(size=len(scores))

        numerator = weighted_strictly_above + (weighted_equal + w_scores) * u

    denominator = total_weight + w_scores
    return numerator / denominator

Weight Estimation

nonconform.weighting

Weight estimation for covariate-shift workflows in weighted conformal inference.

These estimators approximate density ratios w(x) = p_test(x) / p_calibration(x) from calibration and target covariates. Weighted conformal methods can use those ratios under a covariate-shift model, support overlap, and the method's other assumptions. Estimation, misspecification, and clipping can all affect the resulting validity; this module does not infer a distribution-shift guarantee merely by fitting a classifier.

Classes:

Name Description
BaseWeightEstimator

Abstract base class for weight estimators.

IdentityWeightEstimator

Returns uniform weights (no covariate shift).

SklearnWeightEstimator

Wrapper for sklearn probabilistic classifiers.

BootstrapBaggedWeightEstimator

Bootstrap-bagged, batch-specific estimator.

Factory functions

logistic_weight_estimator: Create estimator using Logistic Regression. forest_weight_estimator: Create estimator using Random Forest.

ProbabilisticClassifier

Bases: Protocol

Protocol for classifiers that support probability estimation.

This protocol defines the interface for sklearn-compatible classifiers that can produce probability estimates for weight computation.

fit
fit(X: ndarray, y: ndarray) -> ProbabilisticClassifier

Fit the classifier on training data.

Parameters:

Name Type Description Default
X ndarray

Feature matrix of shape (n_samples, n_features).

required
y ndarray

Target labels of shape (n_samples,).

required

Returns:

Type Description
ProbabilisticClassifier

The fitted classifier instance.

Source code in nonconform/weighting.py
51
52
53
54
55
56
57
58
59
60
61
def fit(self, X: np.ndarray, y: np.ndarray) -> ProbabilisticClassifier:
    """Fit the classifier on training data.

    Args:
        X: Feature matrix of shape (n_samples, n_features).
        y: Target labels of shape (n_samples,).

    Returns:
        The fitted classifier instance.
    """
    ...
predict_proba
predict_proba(X: ndarray) -> np.ndarray

Return probability estimates for samples.

Parameters:

Name Type Description Default
X ndarray

Feature matrix of shape (n_samples, n_features).

required

Returns:

Type Description
ndarray

Probability estimates of shape (n_samples, n_classes).

Source code in nonconform/weighting.py
63
64
65
66
67
68
69
70
71
72
def predict_proba(self, X: np.ndarray) -> np.ndarray:
    """Return probability estimates for samples.

    Args:
        X: Feature matrix of shape (n_samples, n_features).

    Returns:
        Probability estimates of shape (n_samples, n_classes).
    """
    ...

BaseWeightEstimator

Bases: ABC

Abstract base class for weighted conformal density-ratio estimators.

Weight estimators approximate the relative density of target covariates with respect to calibration covariates. A downstream statistical guarantee also requires an appropriate shift model, overlap, and adequate ratios.

Subclasses must implement fit(), _get_stored_weights(), and _score_new_data() to provide specific weight estimation strategies.

fit abstractmethod
fit(
    calibration_samples: ndarray, test_samples: ndarray
) -> None

Estimate density ratio weights.

Source code in nonconform/weighting.py
88
89
90
91
@abstractmethod
def fit(self, calibration_samples: np.ndarray, test_samples: np.ndarray) -> None:
    """Estimate density ratio weights."""
    pass
get_weights
get_weights(
    calibration_samples: ndarray | None = None,
    test_samples: ndarray | None = None,
) -> tuple[np.ndarray, np.ndarray]

Return density ratio weights for calibration and test data.

Parameters:

Name Type Description Default
calibration_samples ndarray | None

Optional calibration data to score. If provided, computes weights for this data using the fitted model. If None, returns stored weights from fit(). Must provide both or neither.

None
test_samples ndarray | None

Optional test data to score. If provided, computes weights for this data using the fitted model. If None, returns stored weights from fit(). Must provide both or neither.

None

Returns:

Type Description
tuple[ndarray, ndarray]

Tuple of (calibration_weights, test_weights) as numpy arrays.

Raises:

Type Description
NotFittedError

If fit() has not been called.

ValueError

If only one of calibration_samples/test_samples is provided.

Source code in nonconform/weighting.py
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
def get_weights(
    self,
    calibration_samples: np.ndarray | None = None,
    test_samples: np.ndarray | None = None,
) -> tuple[np.ndarray, np.ndarray]:
    """Return density ratio weights for calibration and test data.

    Args:
        calibration_samples: Optional calibration data to score. If provided,
            computes weights for this data using the fitted model. If None,
            returns stored weights from fit(). Must provide both or neither.
        test_samples: Optional test data to score. If provided, computes
            weights for this data using the fitted model. If None, returns
            stored weights from fit(). Must provide both or neither.

    Returns:
        Tuple of (calibration_weights, test_weights) as numpy arrays.

    Raises:
        NotFittedError: If fit() has not been called.
        ValueError: If only one of calibration_samples/test_samples is provided.
    """
    if not hasattr(self, "_is_fitted") or not self._is_fitted:
        raise NotFittedError("This weight estimator instance is not fitted yet.")

    if (calibration_samples is None) != (test_samples is None):
        raise ValueError(
            "Must provide both calibration_samples and test_samples, or neither. "
            "Cannot score only one set."
        )

    if calibration_samples is None:
        return self._get_stored_weights()
    else:
        return self._score_new_data(calibration_samples, test_samples)
set_seed
set_seed(seed: int | None) -> None

Set random seed for reproducibility.

Parameters:

Name Type Description Default
seed int | None

Random seed value or None.

required
Source code in nonconform/weighting.py
141
142
143
144
145
146
147
def set_seed(self, seed: int | None) -> None:
    """Set random seed for reproducibility.

    Args:
        seed: Random seed value or None.
    """
    self._seed = seed

IdentityWeightEstimator

IdentityWeightEstimator()

Bases: BaseWeightEstimator

Identity weight estimator that returns uniform weights.

This estimator assumes no covariate shift and returns weights of 1.0 for all samples. It is a baseline that deliberately applies no density-ratio correction.

With Empirical, unit weights reproduce the unweighted p-value formula. Configuring this object as ConformalDetector.weight_estimator still puts the detector in weighted mode, so select() uses weighted conformalized selection rather than ordinary Benjamini-Hochberg.

Source code in nonconform/weighting.py
244
245
246
247
def __init__(self) -> None:
    self._n_calib = 0
    self._n_test = 0
    self._is_fitted = False
fit
fit(
    calibration_samples: ndarray, test_samples: ndarray
) -> None

Fit the identity weight estimator.

Parameters:

Name Type Description Default
calibration_samples ndarray

Array of calibration data samples.

required
test_samples ndarray

Array of test data samples.

required
Source code in nonconform/weighting.py
249
250
251
252
253
254
255
256
257
258
def fit(self, calibration_samples: np.ndarray, test_samples: np.ndarray) -> None:
    """Fit the identity weight estimator.

    Args:
        calibration_samples: Array of calibration data samples.
        test_samples: Array of test data samples.
    """
    self._n_calib = calibration_samples.shape[0]
    self._n_test = test_samples.shape[0]
    self._is_fitted = True

SklearnWeightEstimator

SklearnWeightEstimator(
    base_estimator: ProbabilisticClassifier
    | BaseEstimator
    | None = None,
    clip_quantile: float | None = 0.05,
)

Bases: BaseWeightEstimator

Wrap an sklearn-compatible probabilistic binary classifier.

The configured estimator is cloned before fitting. It must expose fit(), predict_proba(), and classes_ after fitting.

Parameters:

Name Type Description Default
base_estimator ProbabilisticClassifier | BaseEstimator | None

Configured sklearn classifier instance with predict_proba support. Defaults to LogisticRegression.

None
clip_quantile float | None

Quantile for weight clipping (e.g., 0.05 clips to 5th-95th percentile). Use None to disable clipping. Defaults to 0.05.

0.05

Raises:

Type Description
ValueError

If base_estimator does not implement predict_proba.

Examples:

import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

from nonconform.weighting import SklearnWeightEstimator

rng = np.random.default_rng(42)
calibration_samples = rng.normal(size=(120, 3))
test_samples = rng.normal(loc=0.5, size=(80, 3))
estimator = SklearnWeightEstimator(
    base_estimator=make_pipeline(
        StandardScaler(), LogisticRegression(C=1.0, class_weight="balanced")
    )
)
estimator.set_seed(42)
estimator.fit(calibration_samples, test_samples)
calibration_weights, test_weights = estimator.get_weights()
print(calibration_weights.shape, test_weights.shape)
Source code in nonconform/weighting.py
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
def __init__(
    self,
    base_estimator: ProbabilisticClassifier | BaseEstimator | None = None,
    clip_quantile: float | None = 0.05,
) -> None:
    # Default to a sane baseline if nothing is provided
    # Use explicit None check to avoid truthiness evaluation of sklearn estimators
    # (unfitted ensemble estimators raise AttributeError on __len__)
    self.base_estimator = (
        base_estimator
        if base_estimator is not None
        else LogisticRegression(solver="liblinear")
    )
    if clip_quantile is not None and not (0 < clip_quantile < 0.5):
        raise ValueError(
            f"clip_quantile must be in (0, 0.5) or None, got {clip_quantile}."
        )
    self.clip_quantile = clip_quantile

    if not hasattr(self.base_estimator, "predict_proba"):
        raise ValueError(
            f"The provided base_estimator {type(self.base_estimator).__name__} "
            "does not implement 'predict_proba'. Density estimation requires "
            "probability scores. Use SVC(probability=True) or similar."
        )

    # Seed inheritance attribute (may be set by ConformalDetector)
    self._seed: int | None = None

    self.estimator_: ProbabilisticClassifier | None = None
    self._test_class_idx: int | None = None  # Column index for P(Test)
    self._w_calib: np.ndarray | None = None
    self._w_test: np.ndarray | None = None
    self._clip_bounds: tuple[float, float] | None = None
    self._is_fitted = False
fit
fit(
    calibration_samples: ndarray, test_samples: ndarray
) -> None

Fit the weight estimator on calibration and test samples.

Parameters:

Name Type Description Default
calibration_samples ndarray

Array of calibration data samples.

required
test_samples ndarray

Array of test data samples.

required

Raises:

Type Description
ValueError

If calibration_samples is empty.

Source code in nonconform/weighting.py
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
def fit(self, calibration_samples: np.ndarray, test_samples: np.ndarray) -> None:
    """Fit the weight estimator on calibration and test samples.

    Args:
        calibration_samples: Array of calibration data samples.
        test_samples: Array of test data samples.

    Raises:
        ValueError: If calibration_samples is empty.
    """
    if calibration_samples.shape[0] == 0:
        raise ValueError("Calibration samples are empty. Cannot compute weights.")

    # Prepare data (Calib=0, Test=1 labels)
    x_joint, y_joint = self._prepare_training_data(
        calibration_samples, test_samples, self._seed
    )

    self.estimator_ = clone(self.base_estimator)
    if self._seed is not None:
        self._apply_seed_to_estimator(self.estimator_, self._seed)
    self.estimator_.fit(x_joint, y_joint)

    # sklearn sorts classes_ - get correct column index for P(Test)
    self._test_class_idx = int(
        np.where(self.estimator_.classes_ == self.TEST_LABEL)[0][0]
    )

    w_calib, w_test = self._compute_weights(calibration_samples, test_samples)
    self._clip_bounds = self._compute_clip_bounds(
        w_calib, w_test, self.clip_quantile
    )
    self._w_calib, self._w_test = self._clip_weights(
        w_calib, w_test, self._clip_bounds
    )
    self._is_fitted = True

BootstrapBaggedWeightEstimator

BootstrapBaggedWeightEstimator(
    base_estimator: BaseWeightEstimator,
    n_bootstraps: int = 100,
    clip_quantile: float | None = 0.05,
    scoring_mode: Literal["frozen"] = "frozen",
)

Bases: BaseWeightEstimator

Bootstrap-bagged wrapper for weight estimators with instance-wise aggregation.

This estimator repeatedly refits a base weight estimator on balanced bootstrap samples. Geometric averaging can reduce estimator variability in some settings, but it is not universally more accurate or more valid. Compare it against an unbagged estimator on a design representative of deployment.

The algorithm: 1. For each bootstrap iteration: - Resample BOTH sets to balanced sample size (min of calibration and test sizes) - Fit the base estimator on the balanced bootstrap sample - Score ALL original instances using the fitted model (perfect coverage) - Store log(weights) for each instance 2. After all iterations: - Aggregate instance-wise weights using geometric mean (average in log-space) - Optionally clip the aggregated weights at empirical quantiles

Clipping is a numerical stabilization choice that changes the estimated density ratio. Any claimed guarantee must match the clipped construction; a finite observed range alone does not establish the required assumptions.

Seed inheritance

This class uses the _seed attribute pattern for automatic seed inheritance from ConformalDetector.

Parameters:

Name Type Description Default
base_estimator BaseWeightEstimator

Any BaseWeightEstimator instance.

required
n_bootstraps int

Number of bootstrap iterations. Defaults to 100.

100
clip_quantile float | None

Quantile for adaptive clipping. Use None to disable clipping. Defaults to 0.05.

0.05
scoring_mode Literal['frozen']

Weight scoring behavior after fit. Currently only "frozen" is supported, meaning the estimator can only serve the exact calibration/test batches used during fit(). Defaults to "frozen".

'frozen'
Source code in nonconform/weighting.py
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
def __init__(
    self,
    base_estimator: BaseWeightEstimator,
    n_bootstraps: int = 100,
    clip_quantile: float | None = 0.05,
    scoring_mode: Literal["frozen"] = "frozen",
) -> None:
    if n_bootstraps < 1:
        raise ValueError(f"n_bootstraps must be at least 1, got {n_bootstraps}.")
    if clip_quantile is not None and not (0 < clip_quantile < 0.5):
        raise ValueError(f"clip_quantile must be in (0, 0.5), got {clip_quantile}.")
    if scoring_mode != "frozen":
        raise ValueError(
            f"Unsupported scoring_mode {scoring_mode!r}. "
            "BootstrapBaggedWeightEstimator currently supports only "
            "scoring_mode='frozen'."
        )

    self.base_estimator = base_estimator
    self.n_bootstraps = n_bootstraps
    self.clip_quantile = clip_quantile
    self.scoring_mode: Literal["frozen"] = scoring_mode

    # Seed inheritance attribute (set by ConformalDetector)
    self._seed: int | None = None

    self._w_calib: np.ndarray | None = None
    self._w_test: np.ndarray | None = None
    self._calibration_signature: tuple[tuple[int, ...], str, str] | None = None
    self._test_signature: tuple[tuple[int, ...], str, str] | None = None
    self._is_fitted = False
supports_rescoring property
supports_rescoring: bool

Whether this estimator can score arbitrary new batches after fit().

weight_counts property
weight_counts: str

Return diagnostic info about instance-wise weight coverage.

fit
fit(
    calibration_samples: ndarray, test_samples: ndarray
) -> None

Fit the bagged weight estimator with perfect instance coverage.

Parameters:

Name Type Description Default
calibration_samples ndarray

Array of calibration data samples.

required
test_samples ndarray

Array of test data samples.

required

Raises:

Type Description
ValueError

If calibration_samples is empty.

Source code in nonconform/weighting.py
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
def fit(self, calibration_samples: np.ndarray, test_samples: np.ndarray) -> None:
    """Fit the bagged weight estimator with perfect instance coverage.

    Args:
        calibration_samples: Array of calibration data samples.
        test_samples: Array of test data samples.

    Raises:
        ValueError: If calibration_samples is empty.
    """
    if calibration_samples.shape[0] == 0:
        raise ValueError("Calibration samples are empty. Cannot compute weights.")

    n_calib, n_test = len(calibration_samples), len(test_samples)
    sample_size = min(n_calib, n_test)
    rng = np.random.default_rng(self._seed)

    if _bagged_logger.isEnabledFor(logging.INFO):
        _bagged_logger.info(
            f"Bootstrap: n_calib={n_calib}, n_test={n_test}, "
            f"sample_size={sample_size}, n_bootstraps={self.n_bootstraps}. "
            f"Perfect coverage: all instances weighted in all iterations."
        )

    # Online accumulation: sum log-weights (memory efficient)
    sum_log_weights_calib = np.zeros(n_calib)
    sum_log_weights_test = np.zeros(n_test)

    bootstrap_iterator = (
        tqdm(range(self.n_bootstraps), desc="Weighting")
        if _bagged_logger.isEnabledFor(logging.INFO)
        else range(self.n_bootstraps)
    )

    for i in bootstrap_iterator:
        # Resample both sets for balanced comparison
        calib_indices = rng.choice(n_calib, size=sample_size, replace=True)
        test_indices = rng.choice(n_test, size=sample_size, replace=True)
        x_calib_boot = calibration_samples[calib_indices]
        x_test_boot = test_samples[test_indices]

        # Create base estimator with iteration-specific seed
        base_est = deepcopy(self.base_estimator)
        if self._seed is not None:
            derived_seed = derive_seed(i, self._seed)
            if hasattr(base_est, "seed"):
                base_est.seed = derived_seed
            if hasattr(base_est, "_seed"):
                base_est._seed = derived_seed

        # Fit on bootstrap sample, then score ALL original instances
        base_est.fit(x_calib_boot, x_test_boot)
        w_c_all, w_t_all = base_est.get_weights(calibration_samples, test_samples)

        # Accumulate log-weights for geometric mean aggregation
        sum_log_weights_calib += np.log(w_c_all)
        sum_log_weights_test += np.log(w_t_all)

    # Geometric mean aggregation: exp(mean(log-weights))
    w_calib_final = np.exp(sum_log_weights_calib / self.n_bootstraps)
    w_test_final = np.exp(sum_log_weights_test / self.n_bootstraps)

    # Apply clipping after aggregation (use base class static method)
    clip_bounds = BaseWeightEstimator._compute_clip_bounds(
        w_calib_final, w_test_final, self.clip_quantile
    )
    if clip_bounds is None:
        self._w_calib = w_calib_final
        self._w_test = w_test_final
    else:
        clip_min, clip_max = clip_bounds
        self._w_calib = np.clip(w_calib_final, clip_min, clip_max)
        self._w_test = np.clip(w_test_final, clip_min, clip_max)

    self._calibration_signature = self._sample_signature(calibration_samples)
    self._test_signature = self._sample_signature(test_samples)
    self._is_fitted = True

logistic_weight_estimator

logistic_weight_estimator(
    regularization: str | float = "auto",
    clip_quantile: float = 0.05,
    class_weight: str | dict = "balanced",
    max_iter: int = 1000,
) -> SklearnWeightEstimator

Create weight estimator using Logistic Regression.

Note

When used with ConformalDetector, the detector's seed is automatically propagated to the weight estimator for reproducibility.

Parameters:

Name Type Description Default
regularization str | float

Regularization parameter. If 'auto', uses C=1.0. If float, uses as C parameter.

'auto'
clip_quantile float

Quantile for weight clipping. Defaults to 0.05.

0.05
class_weight str | dict

Class weights for LogisticRegression. Defaults to 'balanced'.

'balanced'
max_iter int

Maximum iterations for solver convergence. Defaults to 1000.

1000

Returns:

Type Description
SklearnWeightEstimator

Configured SklearnWeightEstimator instance.

Examples:

import numpy as np

from nonconform import logistic_weight_estimator

rng = np.random.default_rng(42)
calibration_samples = rng.normal(size=(120, 3))
test_samples = rng.normal(loc=0.5, size=(80, 3))
estimator = logistic_weight_estimator(regularization=0.5)
estimator.set_seed(42)
estimator.fit(calibration_samples, test_samples)
calibration_weights, test_weights = estimator.get_weights()
print(calibration_weights.shape, test_weights.shape)
Source code in nonconform/weighting.py
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
def logistic_weight_estimator(
    regularization: str | float = "auto",
    clip_quantile: float = 0.05,
    class_weight: str | dict = "balanced",
    max_iter: int = 1000,
) -> SklearnWeightEstimator:
    """Create weight estimator using Logistic Regression.

    Note:
        When used with ConformalDetector, the detector's seed is automatically
        propagated to the weight estimator for reproducibility.

    Args:
        regularization: Regularization parameter. If 'auto', uses C=1.0.
            If float, uses as C parameter.
        clip_quantile: Quantile for weight clipping. Defaults to 0.05.
        class_weight: Class weights for LogisticRegression. Defaults to 'balanced'.
        max_iter: Maximum iterations for solver convergence. Defaults to 1000.

    Returns:
        Configured SklearnWeightEstimator instance.

    Examples:
        ```python
        import numpy as np

        from nonconform import logistic_weight_estimator

        rng = np.random.default_rng(42)
        calibration_samples = rng.normal(size=(120, 3))
        test_samples = rng.normal(loc=0.5, size=(80, 3))
        estimator = logistic_weight_estimator(regularization=0.5)
        estimator.set_seed(42)
        estimator.fit(calibration_samples, test_samples)
        calibration_weights, test_weights = estimator.get_weights()
        print(calibration_weights.shape, test_weights.shape)
        ```
    """
    from sklearn.pipeline import make_pipeline
    from sklearn.preprocessing import StandardScaler

    c_param = 1.0 if regularization == "auto" else float(regularization)
    base_estimator = make_pipeline(
        StandardScaler(),
        LogisticRegression(
            C=c_param,
            max_iter=max_iter,
            class_weight=class_weight,
        ),
    )
    return SklearnWeightEstimator(
        base_estimator=base_estimator, clip_quantile=clip_quantile
    )

forest_weight_estimator

forest_weight_estimator(
    n_estimators: int = 100,
    max_depth: int | None = 5,
    min_samples_leaf: int = 10,
    clip_quantile: float = 0.05,
) -> SklearnWeightEstimator

Create weight estimator using Random Forest.

Note

When used with ConformalDetector, the detector's seed is automatically propagated to the weight estimator for reproducibility.

Parameters:

Name Type Description Default
n_estimators int

Number of trees in the forest. Defaults to 100.

100
max_depth int | None

Maximum depth of trees. Defaults to 5.

5
min_samples_leaf int

Minimum samples at leaf node. Defaults to 10.

10
clip_quantile float

Quantile for weight clipping. Defaults to 0.05.

0.05

Returns:

Type Description
SklearnWeightEstimator

Configured SklearnWeightEstimator instance.

Examples:

import numpy as np

from nonconform import forest_weight_estimator

rng = np.random.default_rng(42)
calibration_samples = rng.normal(size=(120, 3))
test_samples = rng.normal(loc=0.5, size=(80, 3))
estimator = forest_weight_estimator(n_estimators=50)
estimator.set_seed(42)
estimator.fit(calibration_samples, test_samples)
calibration_weights, test_weights = estimator.get_weights()
print(calibration_weights.shape, test_weights.shape)
Source code in nonconform/weighting.py
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
def forest_weight_estimator(
    n_estimators: int = 100,
    max_depth: int | None = 5,
    min_samples_leaf: int = 10,
    clip_quantile: float = 0.05,
) -> SklearnWeightEstimator:
    """Create weight estimator using Random Forest.

    Note:
        When used with ConformalDetector, the detector's seed is automatically
        propagated to the weight estimator for reproducibility.

    Args:
        n_estimators: Number of trees in the forest. Defaults to 100.
        max_depth: Maximum depth of trees. Defaults to 5.
        min_samples_leaf: Minimum samples at leaf node. Defaults to 10.
        clip_quantile: Quantile for weight clipping. Defaults to 0.05.

    Returns:
        Configured SklearnWeightEstimator instance.

    Examples:
        ```python
        import numpy as np

        from nonconform import forest_weight_estimator

        rng = np.random.default_rng(42)
        calibration_samples = rng.normal(size=(120, 3))
        test_samples = rng.normal(loc=0.5, size=(80, 3))
        estimator = forest_weight_estimator(n_estimators=50)
        estimator.set_seed(42)
        estimator.fit(calibration_samples, test_samples)
        calibration_weights, test_weights = estimator.get_weights()
        print(calibration_weights.shape, test_weights.shape)
        ```
    """
    from sklearn.ensemble import RandomForestClassifier

    base_estimator = RandomForestClassifier(
        n_estimators=n_estimators,
        max_depth=max_depth,
        min_samples_leaf=min_samples_leaf,
        class_weight="balanced",
        n_jobs=-1,
    )
    return SklearnWeightEstimator(
        base_estimator=base_estimator, clip_quantile=clip_quantile
    )

FDR Control

Includes post-hoc FDP bounds, derandomized e-value selection, and weighted low-level expert APIs (weighted_false_discovery_control). For batch workflows, prefer ConformalDetector.select(...). With DerandomizedSplits, it applies e-BH and exposes evidence through last_selection_result; standalone e-value functions remain available for expert use. For simultaneous realized-FDP certification, use detector.fdp_bounds(x, ...) or result.fdp_bounds(...); both return an immutable FDPCertificate. FDPCertificate.from_p_values(...) is the explicit expert array interface. See the FDP migration notes for the intentional replacement of the previous FDP-specific API.

nonconform.fdr

Public false-discovery procedures for conformal anomaly evidence.

Pruning

Bases: Enum

Pruning strategies for weighted FDR control.

Attributes:

Name Type Description
HETEROGENEOUS

Use independent uniform draws for candidate-specific randomized WCS pruning.

HOMOGENEOUS

Use one shared uniform draw for randomized WCS pruning.

DETERMINISTIC

Use the non-randomized WCS pruning rule.

EValueSelectionResult dataclass

EValueSelectionResult(
    e_values: ndarray,
    selected: ndarray,
    alpha: float,
    alpha_bh: float,
    e_threshold: float,
    n_repetitions: int,
    n_calibration: int,
    tie_seed: int | None,
)

Batch e-value FDR selection result.

Attributes:

Name Type Description
e_values ndarray

Uniformly aggregated conformal e-values.

selected ndarray

Boolean e-BH discovery mask.

alpha float

Target FDR level supplied to e-BH.

alpha_bh float

Fixed inner threshold used to construct split e-values.

e_threshold float

Selected e-value cutoff, or infinity when none are selected.

n_repetitions int

Number of split-conformal results aggregated.

n_calibration int

Number of calibration scores in every repetition.

tie_seed int | None

Seed used for randomized score ties, or None when ties were rejected.

FDPCertificate dataclass

FDPCertificate()

Immutable simultaneous certificate for realized FDP at p-value cutoffs.

Construct via detector.fdp_bounds(x), result.fdp_bounds(), or the expert from_p_values() factory. Choose the envelope method before inspecting its curve. Thresholds may then be explored within this fixed testing family. Confidence is simultaneous coverage, not an FDR target.

Evidence and default-grid diagnostics are read-only arrays. Queries never resample. select(t) returns an original-order NumPy mask for p <= t; t is a p-value cutoff, not a requested FDP bound.

Source code in nonconform/fdr.py
74
75
76
77
78
79
def __init__(self) -> None:
    """Require validated construction through the certificate factories."""
    raise TypeError(
        "Use detector.fdp_bounds(), result.fdp_bounds(), or "
        "FDPCertificate.from_p_values()."
    )
p_values property
p_values: ndarray

Read-only evidence in original observation order.

thresholds property
thresholds: ndarray

Read-only default grid of sorted unique observed p-values.

rejection_counts property
rejection_counts: ndarray

Read-only discovery counts on the default grid.

fdp_upper_bounds property
fdp_upper_bounds: ndarray

Read-only FDP upper bounds on the default grid.

precision_lower_bounds property
precision_lower_bounds: ndarray

Read-only precision lower bounds on the default grid.

n_calibration property
n_calibration: int

Calibration sample size.

n_test property
n_test: int

Fixed testing-family size.

confidence property
confidence: float

Simultaneous coverage probability.

method property
method: str

Normalized envelope method.

n_resamples property
n_resamples: int | None

Effective Monte Carlo draws; None for deterministic KS.

seed property
seed: int | None

Monte Carlo seed supplied at construction, or None.

boost property
boost: bool

Whether threshold-specific sharpening is enabled.

lower property
lower: float | None

Effective THC lower truncation; otherwise None.

upper property
upper: float | None

Effective THC upper truncation; otherwise None.

beta property
beta: float | None

Effective THC exponent; otherwise None.

precision property
precision: float | None

Effective BJ inversion tolerance; otherwise None.

from_p_values classmethod
from_p_values(
    p_values: ndarray,
    *,
    n_calibration: int,
    confidence: float = 0.95,
    method: str = "mc_thc",
    n_resamples: int | None = None,
    seed: int | None = None,
    boost: bool = True,
    lower: float | None = None,
    upper: float | None = None,
    beta: float | None = None,
    precision: float | None = None,
) -> FDPCertificate

Certify external p-values; the caller owns provenance assumptions.

Requires unweighted empirical split-conformal p-values from a fixed scoring map and the reference method's exchangeability assumptions. Native detector/snapshot entry points check supported scope; this expert route cannot. Scientific exchangeability is never established by code.

Parameters:

Name Type Description Default
p_values ndarray

Nonempty 1D testing family in [0, 1], in original order.

required
n_calibration int

Positive calibration size shared by all p-values.

required
confidence float

Simultaneous coverage probability in (0, 1).

0.95
method str

mc_thc (default), mc_hc, mc_ks, ks, or mc_bj.

'mc_thc'
n_resamples int | None

Monte Carlo draws; defaults to 1000 for MC methods.

None
seed int | None

Monte Carlo seed only. None draws fresh randomness once.

None
boost bool

Apply threshold-specific sharpening (default True).

True
lower float | None

THC lower truncation, default 0.01.

None
upper float | None

THC upper truncation, default 0.99.

None
beta float | None

THC exponent, default 0.5.

None
precision float | None

BJ inversion tolerance, default 1e-8.

None

Method-specific options must be omitted or None when inapplicable. Deterministic ks accepts neither n_resamples nor seed.

References

Song, Jin, and Candès, "Everywhere Valid Bounds on False Discovery Proportions in Conformal Inference" (2026), arXiv:2605.20726.

Source code in nonconform/fdr.py
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
@classmethod
def from_p_values(
    cls,
    p_values: np.ndarray,
    *,
    n_calibration: int,
    confidence: float = 0.95,
    method: str = "mc_thc",
    n_resamples: int | None = None,
    seed: int | None = None,
    boost: bool = True,
    lower: float | None = None,
    upper: float | None = None,
    beta: float | None = None,
    precision: float | None = None,
) -> FDPCertificate:
    """Certify external p-values; the caller owns provenance assumptions.

    Requires unweighted empirical split-conformal p-values from a fixed
    scoring map and the reference method's exchangeability assumptions.
    Native detector/snapshot entry points check supported scope; this expert
    route cannot. Scientific exchangeability is never established by code.

    Args:
        p_values: Nonempty 1D testing family in [0, 1], in original order.
        n_calibration: Positive calibration size shared by all p-values.
        confidence: Simultaneous coverage probability in (0, 1).
        method: mc_thc (default), mc_hc, mc_ks, ks, or mc_bj.
        n_resamples: Monte Carlo draws; defaults to 1000 for MC methods.
        seed: Monte Carlo seed only. None draws fresh randomness once.
        boost: Apply threshold-specific sharpening (default True).
        lower: THC lower truncation, default 0.01.
        upper: THC upper truncation, default 0.99.
        beta: THC exponent, default 0.5.
        precision: BJ inversion tolerance, default 1e-8.

    Method-specific options must be omitted or None when inapplicable.
    Deterministic ks accepts neither n_resamples nor seed.

    References:
        Song, Jin, and Candès, "Everywhere Valid Bounds on False Discovery
        Proportions in Conformal Inference" (2026), arXiv:2605.20726.
    """
    values = _fdp_bounds.as_p_values("p_values", p_values)
    envelope = _fdp_bounds.prepare_envelope(
        n_calibration=n_calibration,
        n_test=values.size,
        confidence=confidence,
        method=method,
        n_resamples=n_resamples,
        seed=seed,
        boost=boost,
        lower=lower,
        upper=upper,
        beta=beta,
        precision=precision,
    )
    support, counts = np.unique(values, return_counts=True)
    counts = np.cumsum(counts)
    minima = None
    if boost:
        minima = _fdp_bounds.immutable_array(
            np.minimum.accumulate(
                np.minimum(
                    values.size, values.size * envelope.evaluate(support) - counts
                )
            )
        )
    certificate = object.__new__(cls)
    object.__setattr__(
        certificate, "_p_values", _fdp_bounds.immutable_array(values)
    )
    object.__setattr__(
        certificate, "_support", _fdp_bounds.immutable_array(support)
    )
    object.__setattr__(certificate, "_counts", _fdp_bounds.immutable_array(counts))
    object.__setattr__(certificate, "_prefix_minima", minima)
    object.__setattr__(certificate, "_envelope", envelope)
    return certificate
bound_at
bound_at(threshold: float | ndarray) -> float | np.ndarray

Evaluate the simultaneous FDP bound at scalar or vector cutoffs.

Source code in nonconform/fdr.py
193
194
195
196
197
def bound_at(self, threshold: float | np.ndarray) -> float | np.ndarray:
    """Evaluate the simultaneous FDP bound at scalar or vector cutoffs."""
    thresholds, scalar = _fdp_bounds.as_threshold_query(threshold)
    _, bounds = self._query(thresholds)
    return float(bounds[0]) if scalar else bounds
precision_at
precision_at(
    threshold: float | ndarray,
) -> float | np.ndarray

Return 1 - bound_at(threshold), a simultaneous precision lower bound.

Source code in nonconform/fdr.py
199
200
201
def precision_at(self, threshold: float | np.ndarray) -> float | np.ndarray:
    """Return 1 - bound_at(threshold), a simultaneous precision lower bound."""
    return 1.0 - self.bound_at(threshold)
select
select(threshold: float) -> np.ndarray

Return an original-order Boolean NumPy mask for p <= threshold.

Source code in nonconform/fdr.py
203
204
205
206
207
208
def select(self, threshold: float) -> np.ndarray:
    """Return an original-order Boolean NumPy mask for p <= threshold."""
    thresholds, scalar = _fdp_bounds.as_threshold_query(threshold)
    if not scalar:
        raise ValueError("threshold must be a scalar for select().")
    return self._p_values <= thresholds[0]
to_frame
to_frame(thresholds: ndarray | None = None) -> pd.DataFrame

Report a grid; default to sorted unique observed p-values.

Explicit grids preserve order and duplicates and may be empty. Returned tables are independent of the certificate.

Source code in nonconform/fdr.py
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
def to_frame(self, thresholds: np.ndarray | None = None) -> pd.DataFrame:
    """Report a grid; default to sorted unique observed p-values.

    Explicit grids preserve order and duplicates and may be empty.
    Returned tables are independent of the certificate.
    """
    grid = (
        self._support
        if thresholds is None
        else _fdp_bounds.as_thresholds(thresholds)
    )
    counts, bounds = self._query(grid)
    return pd.DataFrame(
        {
            "threshold": grid.copy(),
            "discoveries": counts,
            "fdp_upper_bound": bounds,
            "precision_lower_bound": 1.0 - bounds,
        }
    )

conformal_e_values

conformal_e_values(
    test_scores: ndarray,
    calib_scores: ndarray,
    *,
    alpha_bh: float,
    tie_seed: int | None = None,
) -> np.ndarray

Compute derandomized conformal e-values from split-conformal scores.

This low-level array interface trusts the caller to provide repetitions for the same test family in the same observation order. Repetitions are aggregated uniformly.

Parameters:

Name Type Description Default
test_scores ndarray

Test anomaly scores. Shape (n_test,) for one split or (n_repetitions, n_test) for repeated splits.

required
calib_scores ndarray

Calibration anomaly scores with matching split dimension.

required
alpha_bh float

Inner BH-style threshold for each split construction.

required
tie_seed int | None

None rejects tied scores. A non-negative integer reproducibly randomizes unique secondary ranks for ties.

None

Returns:

Type Description
ndarray

Aggregated e-values of shape (n_test,).

Raises:

Type Description
TypeError

If tie_seed has an unsupported type.

ValueError

If score inputs, alpha_bh, or tie_seed are invalid, or tied scores are found when tie_seed is None.

Source code in nonconform/fdr.py
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
def conformal_e_values(
    test_scores: np.ndarray,
    calib_scores: np.ndarray,
    *,
    alpha_bh: float,
    tie_seed: int | None = None,
) -> np.ndarray:
    """Compute derandomized conformal e-values from split-conformal scores.

    This low-level array interface trusts the caller to provide repetitions for
    the same test family in the same observation order. Repetitions are
    aggregated uniformly.

    Args:
        test_scores: Test anomaly scores. Shape ``(n_test,)`` for one split or
            ``(n_repetitions, n_test)`` for repeated splits.
        calib_scores: Calibration anomaly scores with matching split dimension.
        alpha_bh: Inner BH-style threshold for each split construction.
        tie_seed: ``None`` rejects tied scores. A non-negative integer
            reproducibly randomizes unique secondary ranks for ties.

    Returns:
        Aggregated e-values of shape ``(n_test,)``.

    Raises:
        TypeError: If ``tie_seed`` has an unsupported type.
        ValueError: If score inputs, ``alpha_bh``, or ``tie_seed`` are invalid,
            or tied scores are found when ``tie_seed`` is None.
    """
    return _e_value_core.compute_conformal_e_values(
        test_scores,
        calib_scores,
        alpha_bh=alpha_bh,
        tie_seed=tie_seed,
    )

e_value_false_discovery_control

e_value_false_discovery_control(
    e_values: ndarray, *, alpha: float = 0.05
) -> np.ndarray

Apply the e-BH procedure to non-negative e-values.

Parameters:

Name Type Description Default
e_values ndarray

Non-negative e-values; larger values are stronger evidence.

required
alpha float

Target FDR level in (0, 1).

0.05

Returns:

Type Description
ndarray

Boolean selection mask aligned with e_values.

Source code in nonconform/fdr.py
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
def e_value_false_discovery_control(
    e_values: np.ndarray,
    *,
    alpha: float = 0.05,
) -> np.ndarray:
    """Apply the e-BH procedure to non-negative e-values.

    Args:
        e_values: Non-negative e-values; larger values are stronger evidence.
        alpha: Target FDR level in ``(0, 1)``.

    Returns:
        Boolean selection mask aligned with ``e_values``.
    """
    alpha_value = _fdp_bounds.validate_probability("alpha", alpha)
    e_values_arr = _e_value_core.normalize_e_values(e_values)
    selected, _ = _e_value_core.e_bh_selection(e_values_arr, alpha=alpha_value)
    return selected

select_conformal_e_values

select_conformal_e_values(
    results: Sequence[ConformalResult],
    *,
    alpha: float = 0.05,
    alpha_bh: float | None = None,
    tie_seed: int | None = None,
) -> EValueSelectionResult

Select a fixed test family from repeated split-conformal results.

Native detector provenance checks integrated, unweighted Split results and the recorded test-batch content and ordering. Supply unmodified snapshots: changes to their score arrays are not tracked. Manual or unstamped results are unsupported; expert callers can use :func:conformal_e_values directly.

Parameters:

Name Type Description Default
results Sequence[ConformalResult]

Non-empty sequence of unmodified detector-produced snapshots.

required
alpha float

Target FDR level for the final e-BH procedure.

0.05
alpha_bh float | None

Inner threshold, defaulting to alpha / 10.

None
tie_seed int | None

None rejects ties. A non-negative integer reproducibly randomizes unique secondary ranks for ties.

None

Returns:

Type Description
EValueSelectionResult

Aggregated e-values, final e-BH mask, and diagnostics.

Raises:

Type Description
TypeError

If results or tie_seed has an unsupported type.

ValueError

If provenance, batch identity, score arrays, probabilities, or tied-score handling are unsupported.

Source code in nonconform/fdr.py
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
def select_conformal_e_values(
    results: Sequence[ConformalResult],
    *,
    alpha: float = 0.05,
    alpha_bh: float | None = None,
    tie_seed: int | None = None,
) -> EValueSelectionResult:
    """Select a fixed test family from repeated split-conformal results.

    Native detector provenance checks integrated, unweighted ``Split`` results
    and the recorded test-batch content and ordering. Supply unmodified snapshots:
    changes to their score arrays are not tracked. Manual or unstamped results
    are unsupported; expert callers can use :func:`conformal_e_values` directly.

    Args:
        results: Non-empty sequence of unmodified detector-produced snapshots.
        alpha: Target FDR level for the final e-BH procedure.
        alpha_bh: Inner threshold, defaulting to ``alpha / 10``.
        tie_seed: ``None`` rejects ties. A non-negative integer reproducibly
            randomizes unique secondary ranks for ties.

    Returns:
        Aggregated e-values, final e-BH mask, and diagnostics.

    Raises:
        TypeError: If ``results`` or ``tie_seed`` has an unsupported type.
        ValueError: If provenance, batch identity, score arrays, probabilities,
            or tied-score handling are unsupported.
    """
    alpha_value = _fdp_bounds.validate_probability("alpha", alpha)
    test_scores, calib_scores = _e_value_core.scores_from_results(results)
    return _select_conformal_e_values_from_scores(
        test_scores,
        calib_scores,
        alpha=alpha_value,
        alpha_bh=alpha_bh,
        tie_seed=tie_seed,
    )

weighted_false_discovery_control

weighted_false_discovery_control(
    result: ConformalResult | None,
    *,
    alpha: float = 0.05,
    pruning: Pruning = Pruning.DETERMINISTIC,
    seed: int | None = None,
) -> np.ndarray

Apply weighted conformalized selection to a result bundle.

The result must contain p-values, test and calibration scores, and matching non-negative weights for the same complete testing family. Validity also depends on the weighted-conformal covariate-shift assumptions.

Parameters:

Name Type Description Default
result ConformalResult | None

Weighted detector result for the target family.

required
alpha float

Nominal FDR target in (0, 1).

0.05
pruning Pruning

Deterministic, homogeneous-randomized, or heterogeneous-randomized WCS pruning rule.

DETERMINISTIC
seed int | None

Non-negative seed for randomized pruning, or None.

None

Returns:

Type Description
ndarray

Boolean selection mask aligned with the result's test rows.

Source code in nonconform/fdr.py
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
def weighted_false_discovery_control(
    result: ConformalResult | None,
    *,
    alpha: float = 0.05,
    pruning: Pruning = Pruning.DETERMINISTIC,
    seed: int | None = None,
) -> np.ndarray:
    """Apply weighted conformalized selection to a result bundle.

    The result must contain p-values, test and calibration scores, and matching
    non-negative weights for the same complete testing family. Validity also
    depends on the weighted-conformal covariate-shift assumptions.

    Args:
        result: Weighted detector result for the target family.
        alpha: Nominal FDR target in ``(0, 1)``.
        pruning: Deterministic, homogeneous-randomized, or
            heterogeneous-randomized WCS pruning rule.
        seed: Non-negative seed for randomized pruning, or None.

    Returns:
        Boolean selection mask aligned with the result's test rows.
    """
    p_values, test_scores, calib_scores, test_weights, calib_weights = (
        _wcs.extract_required_fields(result)
    )
    kde_support, use_self_weight = _wcs.extract_kde_support(result)
    return _wcs.run(
        p_values=p_values,
        test_scores=test_scores,
        calib_scores=calib_scores,
        test_weights=test_weights,
        calib_weights=calib_weights,
        alpha=alpha,
        pruning=pruning,
        seed=seed,
        kde_support=kde_support,
        include_self_weight=use_self_weight,
    )

weighted_false_discovery_control_from_arrays

weighted_false_discovery_control_from_arrays(
    *,
    p_values: ndarray,
    test_scores: ndarray,
    calib_scores: ndarray,
    test_weights: ndarray,
    calib_weights: ndarray,
    alpha: float = 0.05,
    pruning: Pruning = Pruning.DETERMINISTIC,
    seed: int | None = None,
) -> np.ndarray

Apply weighted conformalized selection to explicit arrays.

This low-level API cannot verify provenance. All arrays must come from the same calibration construction and complete target family.

Parameters:

Name Type Description Default
p_values ndarray

One p-value per test observation.

required
test_scores ndarray

One anomalous-higher score per test observation.

required
calib_scores ndarray

Calibration scores in the same orientation.

required
test_weights ndarray

Non-negative target-density weights for test observations.

required
calib_weights ndarray

Non-negative target-density weights for calibration observations.

required
alpha float

Nominal FDR target in (0, 1).

0.05
pruning Pruning

WCS pruning rule.

DETERMINISTIC
seed int | None

Non-negative seed for randomized pruning, or None.

None

Returns:

Type Description
ndarray

Boolean selection mask aligned with the test arrays.

Source code in nonconform/fdr.py
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
def weighted_false_discovery_control_from_arrays(
    *,
    p_values: np.ndarray,
    test_scores: np.ndarray,
    calib_scores: np.ndarray,
    test_weights: np.ndarray,
    calib_weights: np.ndarray,
    alpha: float = 0.05,
    pruning: Pruning = Pruning.DETERMINISTIC,
    seed: int | None = None,
) -> np.ndarray:
    """Apply weighted conformalized selection to explicit arrays.

    This low-level API cannot verify provenance. All arrays must come from the
    same calibration construction and complete target family.

    Args:
        p_values: One p-value per test observation.
        test_scores: One anomalous-higher score per test observation.
        calib_scores: Calibration scores in the same orientation.
        test_weights: Non-negative target-density weights for test observations.
        calib_weights: Non-negative target-density weights for calibration
            observations.
        alpha: Nominal FDR target in ``(0, 1)``.
        pruning: WCS pruning rule.
        seed: Non-negative seed for randomized pruning, or None.

    Returns:
        Boolean selection mask aligned with the test arrays.
    """
    return _wcs.run(
        p_values=p_values,
        test_scores=test_scores,
        calib_scores=calib_scores,
        test_weights=test_weights,
        calib_weights=calib_weights,
        alpha=alpha,
        pruning=pruning,
        seed=seed,
    )

Martingales

nonconform.martingales

Exchangeability martingales for sequential conformal evidence.

This module implements p-value-based martingales and alarm statistics for streaming or temporal monitoring workflows. In practice, you feed one conformal p-value at a time and read a running evidence state after each update.

Implemented martingales
  • PowerMartingale
  • SimpleMixtureMartingale
  • SimpleJumperMartingale
  • MixtureMartingale

All classes consume conformal p-values in [0, 1]. Alarm statistics are computed from martingale ratio increments or mixed component statistics and exposed together with the current martingale value in :class:MartingaleState.

AlarmConfig dataclass

AlarmConfig(
    ville_threshold: float | None = None,
    restarted_ville_threshold: float | None = None,
    cusum_threshold: float | None = None,
    shiryaev_roberts_threshold: float | None = None,
)

Optional alarm thresholds for martingale evidence statistics.

Thresholds are disabled when set to None. Each threshold compares against a running statistic in :class:MartingaleState.

ville_threshold and restarted_ville_threshold are Ville thresholds for e-processes. cusum_threshold and shiryaev_roberts_threshold are change-evidence thresholds and should not be interpreted as probability-of-ever-crossing Ville thresholds without a separate theorem for the exact statistic.

Attributes:

Name Type Description
ville_threshold float | None

Threshold for the original cumulative martingale.

restarted_ville_threshold float | None

Threshold for the harmonic restart-mixture e-process.

cusum_threshold float | None

Threshold for the multiplicative CUSUM statistic.

shiryaev_roberts_threshold float | None

Threshold for the Shiryaev-Roberts statistic.

MartingaleState dataclass

MartingaleState(
    step: int,
    p_value: float,
    log_martingale: float,
    martingale: float,
    log_restarted_martingale: float,
    restarted_martingale: float,
    log_cusum: float,
    cusum: float,
    log_shiryaev_roberts: float,
    shiryaev_roberts: float,
    triggered_alarms: tuple[str, ...],
    log_e_value: float = 0.0,
    e_value: float = 1.0,
)

Immutable snapshot of evidence statistics after one update.

Linear-scale values may be 0 or inf after floating-point underflow or overflow; the corresponding log_* field preserves the log-scale state. triggered_alarms reports thresholds crossed at this step and is not a latched alarm history.

e_value is the ordinary capital ratio, not necessarily an increment that reproduces the other statistics (see :class:MixtureMartingale).

BaseMartingale

BaseMartingale(alarm_config: AlarmConfig | None = None)

Bases: ABC

Abstract base class for p-value-driven sequential evidence.

The default update multiplies capital by a non-negative betting factor and updates alarm statistics from that factor. Composite subclasses may instead aggregate component statistics. Ville-threshold interpretation requires the resulting process to be an e-process under the null, which in turn depends on the conditional validity of the input p-values. Merely passing values in [0, 1] does not establish that property.

Source code in nonconform/martingales.py
198
199
200
def __init__(self, alarm_config: AlarmConfig | None = None) -> None:
    self._alarm_config = alarm_config if alarm_config is not None else AlarmConfig()
    self.reset()
state property
state: MartingaleState

Return current state snapshot.

reset
reset() -> None

Reset martingale and alarm statistics to initial values.

Source code in nonconform/martingales.py
207
208
209
210
211
212
213
214
215
216
217
218
def reset(self) -> None:
    """Reset martingale and alarm statistics to initial values."""
    self._step = 0
    self._last_p_value = float("nan")
    self._last_log_increment = 0.0
    self._log_martingale = 0.0
    self._log_active_restarted_mass = float("-inf")
    self._log_restarted_martingale = 0.0
    # CUSUM/SR start at 0 on linear scale -> -inf in log space.
    self._log_cusum = float("-inf")
    self._log_shiryaev_roberts = float("-inf")
    self._reset_method_state()
update_many
update_many(
    p_values: Sequence[float] | ndarray,
) -> list[MartingaleState]

Update state for each p-value in order and return all snapshots.

Source code in nonconform/martingales.py
220
221
222
223
224
def update_many(
    self, p_values: Sequence[float] | np.ndarray
) -> list[MartingaleState]:
    """Update state for each p-value in order and return all snapshots."""
    return [self.update(float(p_value)) for p_value in p_values]
update
update(p_value: float) -> MartingaleState

Ingest one p-value in [0, 1] and return the updated state.

Source code in nonconform/martingales.py
226
227
228
229
230
231
232
233
234
235
236
237
def update(self, p_value: float) -> MartingaleState:
    """Ingest one p-value in ``[0, 1]`` and return the updated state."""
    p_value_validated = _validate_probability(p_value)
    log_increment = self._compute_log_increment(p_value_validated)
    if np.isnan(log_increment):
        raise ValueError("Martingale increment is NaN.")

    self._step += 1
    self._last_p_value = p_value_validated
    self._last_log_increment = log_increment
    self._update_statistics(log_increment)
    return self._current_state()

PowerMartingale

PowerMartingale(
    epsilon: float = 0.5,
    alarm_config: AlarmConfig | None = None,
)

Bases: BaseMartingale

Power martingale with fixed epsilon in (0, 1].

Each p-value contributes the betting factor epsilon * p_value ** (epsilon - 1). Values below one emphasize small p-values; epsilon=1 produces the constant factor one.

Source code in nonconform/martingales.py
321
322
323
324
325
326
327
328
329
def __init__(
    self,
    epsilon: float = 0.5,
    alarm_config: AlarmConfig | None = None,
) -> None:
    self.epsilon = float(epsilon)
    if not (0.0 < self.epsilon <= 1.0):
        raise ValueError(f"epsilon must be in (0, 1], got {self.epsilon}.")
    super().__init__(alarm_config=alarm_config)

SimpleMixtureMartingale

SimpleMixtureMartingale(
    epsilons: Sequence[float] | ndarray | None = None,
    *,
    n_grid: int = 100,
    min_epsilon: float = 0.01,
    alarm_config: AlarmConfig | None = None,
)

Bases: BaseMartingale

Equal-weight mixture of power martingales over an epsilon grid.

When epsilons is omitted, the grid contains n_grid evenly spaced values from min_epsilon through one. The mixture tracks the arithmetic mean of component capitals, evaluated stably in log space.

Its alarm statistics use ratios of that all-history mixture capital. For fixed-weight averages of independently maintained alarm statistics, use :class:MixtureMartingale with separate :class:PowerMartingale experts.

Source code in nonconform/martingales.py
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
def __init__(
    self,
    epsilons: Sequence[float] | np.ndarray | None = None,
    *,
    n_grid: int = 100,
    min_epsilon: float = 0.01,
    alarm_config: AlarmConfig | None = None,
) -> None:
    if epsilons is None:
        if n_grid < 2:
            raise ValueError(f"n_grid must be at least 2, got {n_grid}.")
        if not (0.0 < min_epsilon <= 1.0):
            raise ValueError(f"min_epsilon must be in (0, 1], got {min_epsilon}.")
        self.epsilons = np.linspace(float(min_epsilon), 1.0, int(n_grid))
    else:
        self.epsilons = np.asarray(epsilons, dtype=float)
        if self.epsilons.ndim != 1 or self.epsilons.size == 0:
            raise ValueError("epsilons must be a non-empty 1D sequence.")

    if not np.all(np.isfinite(self.epsilons)):
        raise ValueError("epsilons must be finite.")
    if np.any((self.epsilons <= 0.0) | (self.epsilons > 1.0)):
        raise ValueError("All epsilons must be in (0, 1].")
    self._n_eps = int(self.epsilons.size)
    super().__init__(alarm_config=alarm_config)

MixtureMartingale

MixtureMartingale(
    martingales: Sequence[BaseMartingale],
    *,
    weights: Sequence[float] | None = None,
    alarm_config: AlarmConfig | None = None,
)

Bases: BaseMartingale

Fixed-weight averages of independently updated martingale evidence.

Each component receives the same p-value. Ordinary capital, harmonic restart evidence, CUSUM, and Shiryaev-Roberts statistics are averaged separately; alarm statistics are not computed from mixture capital ratios. Fixed normalized mixtures inherit the corresponding validity guarantees only when every participating component satisfies those guarantees under the common null and filtration. Composition does not remove learning inertia inside an individual expert.

Parameters:

Name Type Description Default
martingales Sequence[BaseMartingale]

Nonempty sequence of reset BaseMartingale instances. Components are deep-copied and exclusively owned by this mixture. Custom martingales and nested mixtures are supported.

required
weights Sequence[float] | None

Finite nonnegative weights with positive total. Normalized once; omitted weights give equal allocation. Zero-weight components are excluded from copying and updates.

None
alarm_config AlarmConfig | None

Thresholds applied to the combined statistics. Component alarms are not propagated.

None
Notes

state.e_value and state.log_e_value describe the ordinary mixture capital ratio only. They cannot reproduce the combined CUSUM, SR, or restart statistics through a shared-increment recursion. For consecutive zero capitals or consecutive infinite log-capitals, the undefined ratio is reported as the neutral diagnostic factor one. Finite log-capitals retain their ratio even when linear values overflow.

A component failure blocks further updates until reset(); the mixture state remains at its last completed update. When used inside an ExchangeabilityMonitor, reset the whole monitor after failure. Aggregation adds O(J) work and storage beyond the J component costs.

Source code in nonconform/martingales.py
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
def __init__(
    self,
    martingales: Sequence[BaseMartingale],
    *,
    weights: Sequence[float] | None = None,
    alarm_config: AlarmConfig | None = None,
) -> None:
    components = tuple(martingales)
    if not components:
        raise ValueError("martingales must be a nonempty sequence.")
    for component in components:
        if not isinstance(component, BaseMartingale):
            raise TypeError("Every component must be a BaseMartingale.")
        if component.state.step != 0:
            raise ValueError("Every component must be reset at construction.")

    try:
        raw_weights = (
            np.ones(len(components), dtype=float)
            if weights is None
            else np.asarray(weights, dtype=float)
        )
    except (TypeError, ValueError, OverflowError) as exc:
        raise ValueError("weights must be a finite numeric sequence.") from exc
    if raw_weights.shape != (len(components),):
        raise ValueError("weights must have one entry per component.")
    if not np.all(np.isfinite(raw_weights)) or np.any(raw_weights < 0):
        raise ValueError("weights must be finite and nonnegative.")
    active = raw_weights > 0
    if not np.any(active):
        raise ValueError("weights must have positive total.")
    # Normalize in log space to preserve tiny positive weights and avoid
    # overflow when several individually finite weights have a huge sum.
    log_weights = np.log(raw_weights[active])
    self._log_weights = log_weights - _logsumexp(log_weights)
    try:
        self._martingales = tuple(
            deepcopy(component)
            for component, participates in zip(components, active, strict=True)
            if participates
        )
    except Exception as exc:
        raise TypeError("martingales must support deep copying.") from exc
    super().__init__(alarm_config=alarm_config)

SimpleJumperMartingale

SimpleJumperMartingale(
    jump: float = 0.01,
    alarm_config: AlarmConfig | None = None,
)

Bases: BaseMartingale

Simple Jumper martingale (Algorithm 1 in Vovk et al.).

This method mixes three betting components with parameters -1, 0, and 1. Before each bet, jump redistributes capital equally among the components; the current p-value then multiplies each component by 1 + epsilon * (p_value - 0.5).

Source code in nonconform/martingales.py
541
542
543
544
545
546
547
548
549
550
def __init__(
    self,
    jump: float = 0.01,
    alarm_config: AlarmConfig | None = None,
) -> None:
    self.jump = float(jump)
    if not (0.0 < self.jump <= 1.0):
        raise ValueError(f"jump must be in (0, 1], got {self.jump}.")
    self._epsilons = np.array([-1.0, 0.0, 1.0], dtype=float)
    super().__init__(alarm_config=alarm_config)

Sequential Monitoring

nonconform.monitoring

Sequential conformal monitoring with exact randomized ranks.

This module supplies the validity-critical first half of an exchangeability martingale workflow: a frozen anomaly scoring rule followed by randomized sequential ranks. The resulting p-values can be consumed by the betting martingales in :mod:nonconform.martingales.

The existing :class:nonconform.ConformalDetector remains the batch and pointwise conformal API. Its fixed-calibration p-values are deliberately not modified by this module.

SequentialRankConformalizer

SequentialRankConformalizer(
    *, tail: Tail = "upper", seed: int | None = None
)

Generate exact randomized sequential conformal p-values.

At each update, the new score is ranked among all scores observed by this object, including the new score itself. Independent uniform tie randomization makes the sequential ranks independent Uniform(0, 1) variables when the complete score sequence is exchangeable. Numeric input validation alone does not establish exchangeability.

Parameters:

Name Type Description Default
tail Tail

"upper" when larger scores are more extreme and "lower" when smaller scores are more extreme. Defaults to "upper".

'upper'
seed int | None

Optional seed for a persistent random number generator.

None
Notes

History grows without a sliding window. The current implementation uses an exact sorted list: rank queries are logarithmic, while insertion is linear in the history length.

Source code in nonconform/monitoring.py
124
125
126
127
128
129
def __init__(self, *, tail: Tail = "upper", seed: int | None = None) -> None:
    if tail not in {"upper", "lower"}:
        raise ValueError("tail must be either 'upper' or 'lower'.")
    self.tail: Tail = tail
    self.seed = _validate_seed(seed)
    self.reset()
count property
count: int

Number of scores currently in the sequential rank history.

scores property
scores: ndarray

Return the sorted score history as a copy.

reset
reset() -> None

Clear score history and restore the initial RNG state.

Source code in nonconform/monitoring.py
141
142
143
144
def reset(self) -> None:
    """Clear score history and restore the initial RNG state."""
    self._sorted_scores: list[float] = []
    self._rng = np.random.default_rng(self.seed)
prime
prime(score: float) -> Self

Add one score to rank history without producing a p-value.

Priming does not update evidence and does not consume a random draw.

Source code in nonconform/monitoring.py
146
147
148
149
150
151
152
def prime(self, score: float) -> Self:
    """Add one score to rank history without producing a p-value.

    Priming does not update evidence and does not consume a random draw.
    """
    insort_right(self._sorted_scores, _validate_score(score))
    return self
prime_many
prime_many(scores: Any) -> Self

Add a one-dimensional score collection without consuming RNG draws.

Source code in nonconform/monitoring.py
154
155
156
157
158
159
160
def prime_many(self, scores: Any) -> Self:
    """Add a one-dimensional score collection without consuming RNG draws."""
    array = _as_1d_scores(scores)
    if array.size:
        sorted_new_scores = np.sort(array).tolist()
        self._sorted_scores = list(merge(self._sorted_scores, sorted_new_scores))
    return self
update
update(score: float) -> float

Insert one score and return its randomized sequential p-value.

The result lies in [0, 1) because the randomized rank uses a uniform draw on a half-open interval.

Source code in nonconform/monitoring.py
162
163
164
165
166
167
168
def update(self, score: float) -> float:
    """Insert one score and return its randomized sequential p-value.

    The result lies in ``[0, 1)`` because the randomized rank uses a uniform
    draw on a half-open interval.
    """
    return self._update_validated(_validate_score(score))
update_many
update_many(scores: Any) -> np.ndarray

Process a one-dimensional score sequence in order.

Source code in nonconform/monitoring.py
188
189
190
191
192
193
def update_many(self, scores: Any) -> np.ndarray:
    """Process a one-dimensional score sequence in order."""
    array = _as_1d_scores(scores)
    return np.asarray(
        [self._update_validated(float(score)) for score in array], dtype=float
    )

MonitorState dataclass

MonitorState(
    rank_step: int,
    score: float,
    martingale_state: MartingaleState,
)

Immutable snapshot from one sequential monitoring update.

evidence_step property
evidence_step: int

Number of observations included in the evidence process.

p_value property
p_value: float

Randomized sequential conformal p-value.

e_value property
e_value: float

Ordinary capital ratio; mixtures aggregate alarm statistics separately.

log_e_value property
log_e_value: float

Natural logarithm of the ordinary capital ratio.

martingale property
martingale: float

Cumulative product-martingale value.

restarted_martingale property
restarted_martingale: float

Harmonic restart-mixture e-process value.

triggered_alarms property
triggered_alarms: tuple[str, ...]

Alarm names whose configured thresholds are currently crossed.

ExchangeabilityMonitor

ExchangeabilityMonitor(
    detector: Any,
    *,
    conformalizer: SequentialRankConformalizer
    | None = None,
    martingale: BaseMartingale | None = None,
    score_polarity: ScorePolarityInput = None,
    seed: int | None = None,
)

Monitor a stream using a frozen scorer and sequential conformal ranks.

The scorer is fitted once and then applied row-wise to reference and stream observations. prime establishes rank history without starting evidence; update generates a randomized sequential p-value and sends it to the configured martingale.

Parameters:

Name Type Description Default
detector Any

PyOD, scikit-learn, or custom anomaly detector.

required
conformalizer SequentialRankConformalizer | None

Stateful sequential rank conformalizer. Defaults to an upper-tail :class:SequentialRankConformalizer.

None
martingale BaseMartingale | None

P-value betting martingale. Defaults to Simple Jumper.

None
score_polarity ScorePolarityInput

Detector score direction, following :class:~nonconform.detector.ConformalDetector semantics.

None
seed int | None

Seed propagated to the scorer and default conformalizer.

None

Supplied conformalizers and martingales are treated as configuration prototypes and deep-copied. The monitor exclusively owns and mutates its scorer, rank history, and evidence process.

Notes

Sequential-rank validity requires the priming and monitored scores to be exchangeable under the null, conditional on a fixed training-only scoring construction. Ville guarantees additionally require a valid e-process. Do not refit the scorer during an episode.

Source code in nonconform/monitoring.py
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
def __init__(
    self,
    detector: Any,
    *,
    conformalizer: SequentialRankConformalizer | None = None,
    martingale: BaseMartingale | None = None,
    score_polarity: ScorePolarityInput = None,
    seed: int | None = None,
) -> None:
    self.seed = _validate_seed(seed)
    adapted_detector = adapt(detector)
    if score_polarity is None:
        resolved_polarity = resolve_implicit_score_polarity(adapted_detector)
    else:
        resolved_polarity = resolve_score_polarity(adapted_detector, score_polarity)
    normalized_detector = apply_score_polarity(adapted_detector, resolved_polarity)
    owned_detector: AnomalyDetector = set_params(
        _copy_component("detector", normalized_detector), self.seed
    )
    self._initialize_owned(
        detector=owned_detector,
        conformalizer=conformalizer,
        martingale=martingale,
        seed=self.seed,
        is_fitted=False,
        n_features_in=None,
    )
is_fitted property
is_fitted: bool

Whether the scoring rule is fitted and frozen.

state property
state: MonitorState | None

Return the latest monitoring state, or None before the first update.

from_split_detector classmethod
from_split_detector(
    detector: ConformalDetector,
    *,
    conformalizer: SequentialRankConformalizer
    | None = None,
    martingale: BaseMartingale | None = None,
    seed: int | None = None,
) -> ExchangeabilityMonitor

Create a monitor from a fitted unweighted Split detector.

The fitted scoring model is copied and frozen. Existing calibration scores initialize sequential rank history; they are not reused as a fixed empirical CDF and do not contribute to martingale capital.

Parameters:

Name Type Description Default
detector ConformalDetector

Fitted, unweighted detector using Split and a single retained scoring model.

required
conformalizer SequentialRankConformalizer | None

Optional empty conformalizer configuration prototype.

None
martingale BaseMartingale | None

Optional reset martingale configuration prototype.

None
seed int | None

Seed for the monitor's conformalizer, or None.

None

Returns:

Type Description
ExchangeabilityMonitor

A fitted monitor primed with copied calibration scores.

Source code in nonconform/monitoring.py
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
@classmethod
def from_split_detector(
    cls,
    detector: ConformalDetector,
    *,
    conformalizer: SequentialRankConformalizer | None = None,
    martingale: BaseMartingale | None = None,
    seed: int | None = None,
) -> ExchangeabilityMonitor:
    """Create a monitor from a fitted unweighted ``Split`` detector.

    The fitted scoring model is copied and frozen. Existing calibration
    scores initialize sequential rank history; they are not reused as a
    fixed empirical CDF and do not contribute to martingale capital.

    Args:
        detector: Fitted, unweighted detector using ``Split`` and a single
            retained scoring model.
        conformalizer: Optional empty conformalizer configuration prototype.
        martingale: Optional reset martingale configuration prototype.
        seed: Seed for the monitor's conformalizer, or None.

    Returns:
        A fitted monitor primed with copied calibration scores.
    """
    snapshot = _snapshot_split_detector(detector)
    monitor = cls.__new__(cls)
    monitor._initialize_owned(
        detector=snapshot.detector,
        conformalizer=conformalizer,
        martingale=martingale,
        seed=seed,
        is_fitted=True,
        n_features_in=snapshot.n_features_in,
    )
    monitor._conformalizer.prime_many(snapshot.calibration_scores)
    return monitor
fit
fit(x: DataFrame | ndarray) -> Self

Fit the scoring rule once on a proper training set.

Parameters:

Name Type Description Default
x DataFrame | ndarray

Finite two-dimensional fitting data. It must be independent of the later null rank sequence for the documented validity scope.

required

Returns:

Type Description
Self

This monitor with empty rank and evidence state.

Source code in nonconform/monitoring.py
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
def fit(self, x: pd.DataFrame | np.ndarray) -> Self:
    """Fit the scoring rule once on a proper training set.

    Args:
        x: Finite two-dimensional fitting data. It must be independent of
            the later null rank sequence for the documented validity scope.

    Returns:
        This monitor with empty rank and evidence state.
    """
    x_array = _as_2d_array("x", x)
    self._detector.fit(x_array)
    self._is_fitted = True
    self._n_features_in = int(x_array.shape[1])
    self.reset()
    return self
reset
reset() -> None

Reset rank and evidence state while retaining the fitted scorer.

A reset starts a new monitoring episode. Repeated episodes need their own error-budget accounting if a lifetime false-alarm guarantee is required.

Source code in nonconform/monitoring.py
399
400
401
402
403
404
405
406
407
408
def reset(self) -> None:
    """Reset rank and evidence state while retaining the fitted scorer.

    A reset starts a new monitoring episode.  Repeated episodes need their
    own error-budget accounting if a lifetime false-alarm guarantee is
    required.
    """
    self._conformalizer.reset()
    self._martingale.reset()
    self._last_state = None
prime
prime(x: DataFrame | ndarray) -> Self

Score reference observations and add them only to rank history.

Parameters:

Name Type Description Default
x DataFrame | ndarray

Finite two-dimensional reference batch with the fitted feature count.

required

Returns:

Type Description
Self

This monitor after extending its rank history.

Source code in nonconform/monitoring.py
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
def prime(self, x: pd.DataFrame | np.ndarray) -> Self:
    """Score reference observations and add them only to rank history.

    Args:
        x: Finite two-dimensional reference batch with the fitted feature
            count.

    Returns:
        This monitor after extending its rank history.
    """
    self._require_fitted()
    if self._martingale.state.step != 0:
        raise RuntimeError(
            "prime() is unavailable after evidence monitoring starts."
        )
    x_array = self._validate_feature_batch("x", x)
    scores = np.asarray([self._score_one(row) for row in x_array], dtype=float)
    self._conformalizer.prime_many(scores)
    return self
update
update(x: Series | ndarray) -> MonitorState

Score one observation and update sequential evidence.

Parameters:

Name Type Description Default
x Series | ndarray

One finite feature vector with the fitted feature count.

required

Returns:

Type Description
MonitorState

The immutable state after ranking, betting, and alarm evaluation.

Source code in nonconform/monitoring.py
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
def update(self, x: pd.Series | np.ndarray) -> MonitorState:
    """Score one observation and update sequential evidence.

    Args:
        x: One finite feature vector with the fitted feature count.

    Returns:
        The immutable state after ranking, betting, and alarm evaluation.
    """
    self._require_fitted()
    try:
        x_array = np.asarray(x, dtype=float)
    except (TypeError, ValueError, OverflowError) as exc:
        raise ValueError("x must be a finite numeric feature vector.") from exc
    if x_array.ndim != 1:
        raise ValueError("x must be a one-dimensional feature vector.")
    self._validate_feature_count(x_array)
    score = self._score_one(x_array)
    p_value = self._conformalizer.update(score)
    martingale_state = self._martingale.update(p_value)
    state = MonitorState(
        rank_step=self._conformalizer.count,
        score=score,
        martingale_state=martingale_state,
    )
    self._last_state = state
    return state
update_many
update_many(x: DataFrame | ndarray) -> list[MonitorState]

Update sequentially for every row in input order.

The input is accepted as a two-dimensional batch for convenience, but the method preserves row order and performs one sequential update per row.

Source code in nonconform/monitoring.py
458
459
460
461
462
463
464
465
466
467
def update_many(self, x: pd.DataFrame | np.ndarray) -> list[MonitorState]:
    """Update sequentially for every row in input order.

    The input is accepted as a two-dimensional batch for convenience, but
    the method preserves row order and performs one sequential update per
    row.
    """
    self._require_fitted()
    x_array = self._validate_feature_batch("x", x)
    return [self.update(row) for row in x_array]

Metrics

nonconform.metrics

Public helpers for score aggregation and labeled evaluation.

false_discovery_rate retains its v1 public name but returns the realized false discovery proportion for one supplied testing family. statistical_power similarly returns the realized true positive rate. Expected FDR and statistical power are repeated-sampling properties, not quantities identified by one labeled family.

aggregate

aggregate(method: str, scores: ndarray) -> np.ndarray

Aggregate anomaly scores using a specified method.

Applies a chosen aggregation technique to a 2D array of anomaly scores, where each row represents scores from a different model and each column corresponds to a data sample.

Parameters:

Name Type Description Default
method str

The aggregation method to apply.

required
scores ndarray

A 2D array of anomaly scores. Rows = different models, columns = data samples. Aggregation is performed along axis=0.

required

Returns:

Type Description
ndarray

Array of aggregated anomaly scores with length equal to number

ndarray

of columns in input.

Raises:

Type Description
ValueError

If the method is not a supported aggregation type.

Examples:

>>> import numpy as np
>>> from nonconform.metrics import aggregate
>>> scores = np.array([[1, 2, 3], [4, 5, 6]])
>>> aggregate("mean", scores)
array([2.5, 3.5, 4.5])
Source code in nonconform/_internal/math_utils.py
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
def aggregate(method: str, scores: np.ndarray) -> np.ndarray:
    """Aggregate anomaly scores using a specified method.

    Applies a chosen aggregation technique to a 2D array of anomaly scores,
    where each row represents scores from a different model and each column
    corresponds to a data sample.

    Args:
        method: The aggregation method to apply.
        scores: A 2D array of anomaly scores. Rows = different models,
            columns = data samples. Aggregation is performed along axis=0.

    Returns:
        Array of aggregated anomaly scores with length equal to number
        of columns in input.

    Raises:
        ValueError: If the method is not a supported aggregation type.

    Examples:
        >>> import numpy as np
        >>> from nonconform.metrics import aggregate
        >>> scores = np.array([[1, 2, 3], [4, 5, 6]])
        >>> aggregate("mean", scores)
        array([2.5, 3.5, 4.5])
    """
    normalized = normalize_aggregation_method(method)
    match normalized:
        case "mean":
            return np.mean(scores, axis=0)
        case "median":
            return np.median(scores, axis=0)
        case "minimum":
            return np.min(scores, axis=0)
        case "maximum":
            return np.max(scores, axis=0)
    assert_never(normalized)

false_discovery_rate

false_discovery_rate(y: ndarray, y_hat: ndarray) -> float

Calculate the realized false discovery proportion for one labeled family.

The returned quantity is FP / (FP + TP) for the supplied realization. Its expectation over repeated testing families is the false discovery rate (FDR). The function name is retained for public API compatibility.

If there are no predicted positives, this function returns a realized FDP of 0.0.

Parameters:

Name Type Description Default
y ndarray

True binary labels (1 = positive/anomaly, 0 = negative/normal).

required
y_hat ndarray

Predicted binary labels.

required

Returns:

Type Description
float

The realized false discovery proportion.

Examples:

>>> import numpy as np
>>> from nonconform.metrics import false_discovery_rate
>>> y = np.array([1, 0, 1, 0])
>>> y_hat = np.array([1, 1, 0, 0])  # 1 TP, 1 FP
>>> float(false_discovery_rate(y, y_hat))
0.5
Source code in nonconform/_internal/math_utils.py
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
def false_discovery_rate(y: np.ndarray, y_hat: np.ndarray) -> float:
    """Calculate the realized false discovery proportion for one labeled family.

    The returned quantity is ``FP / (FP + TP)`` for the supplied realization.
    Its expectation over repeated testing families is the false discovery rate
    (FDR). The function name is retained for public API compatibility.

    If there are no predicted positives, this function returns a realized FDP
    of 0.0.

    Args:
        y: True binary labels (1 = positive/anomaly, 0 = negative/normal).
        y_hat: Predicted binary labels.

    Returns:
        The realized false discovery proportion.

    Examples:
        >>> import numpy as np
        >>> from nonconform.metrics import false_discovery_rate
        >>> y = np.array([1, 0, 1, 0])
        >>> y_hat = np.array([1, 1, 0, 0])  # 1 TP, 1 FP
        >>> float(false_discovery_rate(y, y_hat))
        0.5
    """
    y_true = y.astype(bool)
    y_pred = y_hat.astype(bool)

    true_positives = np.sum(y_pred & y_true)
    false_positives = np.sum(y_pred & ~y_true)

    total_predicted_positives = true_positives + false_positives

    if total_predicted_positives == 0:
        return 0.0

    return false_positives / total_predicted_positives

statistical_power

statistical_power(y: ndarray, y_hat: ndarray) -> float

Calculate realized recall (true positive rate) for one labeled family.

The returned quantity is TP / (TP + FN) for the supplied realization. Its expectation under a specified data-generating process is statistical power. The function name is retained for public API compatibility.

If there are no actual positives, this function returns a realized true positive rate of 0.0.

Parameters:

Name Type Description Default
y ndarray

True binary labels (1 = positive/anomaly, 0 = negative/normal).

required
y_hat ndarray

Predicted binary labels.

required

Returns:

Type Description
float

The realized true positive rate.

Examples:

>>> import numpy as np
>>> from nonconform.metrics import statistical_power
>>> y = np.array([1, 0, 1, 0])
>>> y_hat = np.array([1, 1, 0, 0])  # 1 TP, 1 FN
>>> float(statistical_power(y, y_hat))
0.5
Source code in nonconform/_internal/math_utils.py
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
def statistical_power(y: np.ndarray, y_hat: np.ndarray) -> float:
    """Calculate realized recall (true positive rate) for one labeled family.

    The returned quantity is ``TP / (TP + FN)`` for the supplied realization.
    Its expectation under a specified data-generating process is statistical
    power. The function name is retained for public API compatibility.

    If there are no actual positives, this function returns a realized true
    positive rate of 0.0.

    Args:
        y: True binary labels (1 = positive/anomaly, 0 = negative/normal).
        y_hat: Predicted binary labels.

    Returns:
        The realized true positive rate.

    Examples:
        >>> import numpy as np
        >>> from nonconform.metrics import statistical_power
        >>> y = np.array([1, 0, 1, 0])
        >>> y_hat = np.array([1, 1, 0, 0])  # 1 TP, 1 FN
        >>> float(statistical_power(y, y_hat))
        0.5
    """
    y_bool = y.astype(bool)
    y_hat_bool = y_hat.astype(bool)

    true_positives = np.sum(y_bool & y_hat_bool)
    false_negatives = np.sum(y_bool & ~y_hat_bool)
    total_actual_positives = true_positives + false_negatives

    if total_actual_positives == 0:
        return 0.0

    return true_positives / total_actual_positives

Enumerations

nonconform.enums

Public enumeration constants for nonconform.

ConformalMode

Bases: Enum

Model retention modes for conformal resampling strategies.

Attributes:

Name Type Description
PLUS

Retain calibration-time models and aggregate their test scores.

SINGLE_MODEL

Fit and retain one final model after constructing calibration scores.

Distribution

Bases: Enum

Reserved distribution choices for randomized size configurations.

The enum remains part of the v1 public surface through nonconform.enums. Current public calibration strategy constructors do not consume it.

Attributes:

Name Type Description
BETA_BINOMIAL

Beta-binomial distribution for drawing validation fractions.

UNIFORM

Discrete uniform distribution over a specified range.

GRID

Discrete distribution over a specified set of values.

Kernel

Bases: Enum

Kernel functions for KDE-based score-tail estimation.

Attributes:

Name Type Description
GAUSSIAN

Gaussian (normal) kernel.

EXPONENTIAL

Exponential kernel.

BOX

Box (uniform) kernel.

TRIANGULAR

Triangular kernel.

EPANECHNIKOV

Epanechnikov kernel.

BIWEIGHT

Biweight (quartic) kernel.

TRIWEIGHT

Triweight kernel.

TRICUBE

Tricube kernel.

COSINE

Cosine kernel.

Pruning

Bases: Enum

Pruning strategies for weighted FDR control.

Attributes:

Name Type Description
HETEROGENEOUS

Use independent uniform draws for candidate-specific randomized WCS pruning.

HOMOGENEOUS

Use one shared uniform draw for randomized WCS pruning.

DETERMINISTIC

Use the non-randomized WCS pruning rule.

ScorePolarity

Bases: Enum

Score direction conventions for anomaly detectors.

Attributes:

Name Type Description
AUTO

Strictly infer polarity from recognized detector families and raise for an unknown custom detector.

HIGHER_IS_ANOMALOUS

Higher scores indicate more anomalous samples.

HIGHER_IS_NORMAL

Higher scores indicate more normal samples.

TieBreakMode

Bases: Enum

Tie-breaking modes for empirical p-value estimation.

Attributes:

Name Type Description
CLASSICAL

Deterministic empirical conformal formula.

RANDOMIZED

Randomized smoothing with uniform tie-breaking.

Data Structures

nonconform.structures

Core data structures and protocols for nonconform.

This module provides the fundamental types used throughout the package:

Classes:

Name Description
AnomalyDetector

Protocol defining the detector interface.

ConformalResult

Container for conformal inference outputs.

AnomalyDetector

Bases: Protocol

Protocol defining the interface for anomaly detectors.

A supported PyOD model, a recognized scikit-learn estimator, or a custom object can be used with nonconform when it provides this interface. The object must also support shallow and deep copying because resampling strategies create independent detector replicas.

Required methods

fit: Train the detector on data decision_function: Compute anomaly scores get_params: Retrieve detector parameters set_params: Configure detector parameters

Examples:

from sklearn.ensemble import IsolationForest

from nonconform.structures import AnomalyDetector

detector: AnomalyDetector = IsolationForest(random_state=42)
print(isinstance(detector, AnomalyDetector))
fit
fit(X: ndarray, y: ndarray | None = None) -> Self

Train the anomaly detector.

Parameters:

Name Type Description Default
X ndarray

Training data of shape (n_samples, n_features).

required
y ndarray | None

Ignored. Present for API consistency.

None

Returns:

Type Description
Self

The fitted detector instance.

Source code in nonconform/structures.py
56
57
58
59
60
61
62
63
64
65
66
def fit(self, X: np.ndarray, y: np.ndarray | None = None) -> Self:
    """Train the anomaly detector.

    Args:
        X: Training data of shape (n_samples, n_features).
        y: Ignored. Present for API consistency.

    Returns:
        The fitted detector instance.
    """
    ...
decision_function
decision_function(X: ndarray) -> np.ndarray

Compute anomaly scores for samples.

Score direction is detector-specific. Pass the corresponding score_polarity to :class:~nonconform.detector.ConformalDetector when it cannot be inferred safely.

Parameters:

Name Type Description Default
X ndarray

Data of shape (n_samples, n_features).

required

Returns:

Type Description
ndarray

Anomaly scores of shape (n_samples,).

Source code in nonconform/structures.py
68
69
70
71
72
73
74
75
76
77
78
79
80
81
def decision_function(self, X: np.ndarray) -> np.ndarray:
    """Compute anomaly scores for samples.

    Score direction is detector-specific. Pass the corresponding
    ``score_polarity`` to :class:`~nonconform.detector.ConformalDetector`
    when it cannot be inferred safely.

    Args:
        X: Data of shape (n_samples, n_features).

    Returns:
        Anomaly scores of shape (n_samples,).
    """
    ...
get_params
get_params(deep: bool = True) -> dict[str, Any]

Get parameters for this detector.

Parameters:

Name Type Description Default
deep bool

If True, return parameters for sub-objects.

True

Returns:

Type Description
dict[str, Any]

Parameter names mapped to their values.

Source code in nonconform/structures.py
83
84
85
86
87
88
89
90
91
92
def get_params(self, deep: bool = True) -> dict[str, Any]:
    """Get parameters for this detector.

    Args:
        deep: If True, return parameters for sub-objects.

    Returns:
        Parameter names mapped to their values.
    """
    ...
set_params
set_params(**params: Any) -> Self

Set parameters for this detector.

Parameters:

Name Type Description Default
**params Any

Detector parameters.

{}

Returns:

Type Description
Self

The detector instance.

Source code in nonconform/structures.py
 94
 95
 96
 97
 98
 99
100
101
102
103
def set_params(self, **params: Any) -> Self:
    """Set parameters for this detector.

    Args:
        **params: Detector parameters.

    Returns:
        The detector instance.
    """
    ...

ConformalResult dataclass

ConformalResult(
    p_values: ndarray | None = None,
    test_scores: ndarray | None = None,
    calib_scores: ndarray | None = None,
    test_weights: ndarray | None = None,
    calib_weights: ndarray | None = None,
    metadata: dict[str, Any] = dict(),
)

Snapshot of detector outputs for downstream procedures.

This dataclass holds the latest p-values or score-tail estimates, raw scores, and optional weights produced by a detector call.

Attributes:

Name Type Description
p_values ndarray | None

Values produced by the configured estimation strategy, or None when only scores were requested. With Empirical, these are rank-based conformal p-values.

test_scores ndarray | None

Aggregated, anomalous-higher scores for test instances.

calib_scores ndarray | None

Anomalous-higher scores for the calibration set, or None for aggregated raw DerandomizedSplits scores, which have no single corresponding calibration distribution.

test_weights ndarray | None

Importance weights for test instances (weighted mode only).

calib_weights ndarray | None

Importance weights for calibration instances.

metadata dict[str, Any]

Method metadata, including the strategy, estimator, and weighted-mode marker for p-value computations.

Examples:

import numpy as np

from nonconform.structures import ConformalResult

result = ConformalResult(
    p_values=np.array([0.50, 0.02]),
    test_scores=np.array([0.1, 2.4]),
    calib_scores=np.array([-0.2, 0.0, 0.3, 0.8]),
    metadata={"nonconform": {"weighted": False}},
)
print(result.p_values)
print(result.metadata)
fdp_bounds
fdp_bounds(
    *,
    confidence: float = 0.95,
    method: str = "mc_thc",
    n_resamples: int | None = None,
    seed: int | None = None,
    boost: bool = True,
    lower: float | None = None,
    upper: float | None = None,
    beta: float | None = None,
    precision: float | None = None,
) -> FDPCertificate

Certify this snapshot without rescoring.

Requires an unmodified native snapshot of unweighted empirical Split inference, including detached calibration. Scope and batch dimensions are checked; these checks do not prove array integrity or scientific exchangeability. For external p-values use FDPCertificate.from_p_values.

Options and defaults match :meth:nonconform.fdr.FDPCertificate.from_p_values. Confidence is simultaneous coverage, not an FDR target. The Monte Carlo seed is independent of the detector seed. Query cutoffs on the returned immutable certificate; later edits to this snapshot cannot affect it.

Source code in nonconform/structures.py
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
def fdp_bounds(
    self,
    *,
    confidence: float = 0.95,
    method: str = "mc_thc",
    n_resamples: int | None = None,
    seed: int | None = None,
    boost: bool = True,
    lower: float | None = None,
    upper: float | None = None,
    beta: float | None = None,
    precision: float | None = None,
) -> FDPCertificate:
    """Certify this snapshot without rescoring.

    Requires an unmodified native snapshot of unweighted empirical Split
    inference, including detached calibration. Scope and batch dimensions
    are checked; these checks do not prove array integrity or scientific
    exchangeability. For external p-values use FDPCertificate.from_p_values.

    Options and defaults match :meth:`nonconform.fdr.FDPCertificate.from_p_values`.
    Confidence is simultaneous coverage, not an FDR target. The Monte Carlo
    seed is independent of the detector seed. Query cutoffs on the returned
    immutable certificate; later edits to this snapshot cannot affect it.
    """
    from nonconform._internal.fdp_bounds import validate_result_scope
    from nonconform.fdr import FDPCertificate

    n_calibration = validate_result_scope(self)
    return FDPCertificate.from_p_values(
        self.p_values,
        n_calibration=n_calibration,
        confidence=confidence,
        method=method,
        n_resamples=n_resamples,
        seed=seed,
        boost=boost,
        lower=lower,
        upper=upper,
        beta=beta,
        precision=precision,
    )
copy
copy() -> ConformalResult

Return a copy with arrays and metadata fully duplicated.

Returns:

Type Description
ConformalResult

A new ConformalResult with copied arrays and deep-copied metadata.

Source code in nonconform/structures.py
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
def copy(self) -> ConformalResult:
    """Return a copy with arrays and metadata fully duplicated.

    Returns:
        A new ConformalResult with copied arrays and deep-copied metadata.
    """

    def _copy_arr(arr: np.ndarray | None) -> np.ndarray | None:
        return arr.copy() if arr is not None else None

    copied = ConformalResult(
        p_values=_copy_arr(self.p_values),
        test_scores=_copy_arr(self.test_scores),
        calib_scores=_copy_arr(self.calib_scores),
        test_weights=_copy_arr(self.test_weights),
        calib_weights=_copy_arr(self.calib_weights),
        metadata=deepcopy(self.metadata),
    )
    copied._provenance = self._provenance
    return copied

Adapters

nonconform.adapters

External detector adapters for nonconform.

ScorePolarityAdapter

ScorePolarityAdapter(
    detector: AnomalyDetector, score_polarity: ScorePolarity
)

Normalize a wrapped detector to anomalous-higher scores.

Parameters:

Name Type Description Default
detector AnomalyDetector

Detector implementing :class:AnomalyDetector.

required
score_polarity ScorePolarity

Convention of its raw scores. Automatic polarity is not accepted here.

required
Source code in nonconform/adapters.py
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
def __init__(
    self,
    detector: AnomalyDetector,
    score_polarity: ScorePolarity,
) -> None:
    if score_polarity not in {
        ScorePolarity.HIGHER_IS_ANOMALOUS,
        ScorePolarity.HIGHER_IS_NORMAL,
    }:
        raise ValueError(
            "ScorePolarityAdapter requires explicit non-auto score polarity."
        )
    self._detector = detector
    self._score_polarity = score_polarity
    self._multiplier = (
        1.0 if score_polarity is ScorePolarity.HIGHER_IS_ANOMALOUS else -1.0
    )
fit
fit(X: ndarray, y: ndarray | None = None) -> Self

Fit wrapped detector.

Source code in nonconform/adapters.py
295
296
297
298
def fit(self, X: np.ndarray, y: np.ndarray | None = None) -> Self:
    """Fit wrapped detector."""
    self._detector.fit(X, y)
    return self
decision_function
decision_function(X: ndarray) -> np.ndarray

Return scores transformed to anomalous-higher convention.

Source code in nonconform/adapters.py
300
301
302
303
def decision_function(self, X: np.ndarray) -> np.ndarray:
    """Return scores transformed to anomalous-higher convention."""
    scores = np.asarray(self._detector.decision_function(X), dtype=float)
    return self._multiplier * scores
get_params
get_params(deep: bool = True) -> dict[str, Any]

Delegate parameter retrieval to wrapped detector.

Source code in nonconform/adapters.py
305
306
307
def get_params(self, deep: bool = True) -> dict[str, Any]:
    """Delegate parameter retrieval to wrapped detector."""
    return self._detector.get_params(deep=deep)
set_params
set_params(**params: Any) -> Self

Delegate parameter updates to wrapped detector.

Source code in nonconform/adapters.py
309
310
311
312
def set_params(self, **params: Any) -> Self:
    """Delegate parameter updates to wrapped detector."""
    self._detector.set_params(**params)
    return self

PyODAdapter

PyODAdapter(detector: Any)

Wrap a PyOD detector with the public anomaly-detector protocol.

Parameters:

Name Type Description Default
detector Any

Configured PyOD detector.

required

Raises:

Type Description
ImportError

If PyOD is not installed.

Source code in nonconform/adapters.py
353
354
355
356
357
def __init__(self, detector: Any) -> None:
    """Initialize adapter for a PyOD detector."""
    if not PYOD_AVAILABLE:
        raise ImportError("PyOD is not installed. Install with: pip install pyod")
    self._detector = detector
fit
fit(X: ndarray, y: ndarray | None = None) -> Self

Fit wrapped detector.

Source code in nonconform/adapters.py
359
360
361
362
def fit(self, X: np.ndarray, y: np.ndarray | None = None) -> Self:
    """Fit wrapped detector."""
    self._detector.fit(X, y)
    return self
decision_function
decision_function(X: ndarray) -> np.ndarray

Return anomaly scores from wrapped detector.

Source code in nonconform/adapters.py
364
365
366
def decision_function(self, X: np.ndarray) -> np.ndarray:
    """Return anomaly scores from wrapped detector."""
    return self._detector.decision_function(X)
get_params
get_params(deep: bool = True) -> dict[str, Any]

Delegate parameter retrieval to wrapped detector.

Source code in nonconform/adapters.py
368
369
370
def get_params(self, deep: bool = True) -> dict[str, Any]:
    """Delegate parameter retrieval to wrapped detector."""
    return self._detector.get_params(deep=deep)
set_params
set_params(**params: Any) -> Self

Delegate parameter updates to wrapped detector.

Source code in nonconform/adapters.py
372
373
374
375
def set_params(self, **params: Any) -> Self:
    """Delegate parameter updates to wrapped detector."""
    self._detector.set_params(**params)
    return self

adapt

adapt(detector: Any) -> AnomalyDetector

Return a detector that satisfies the public anomaly-detector protocol.

Parameters:

Name Type Description Default
detector Any

PyOD, scikit-learn, or custom detector object.

required

Returns:

Type Description
AnomalyDetector

The original structurally compatible object or a PyOD adapter.

Raises:

Type Description
ValueError

If the detector is a blocked batch-adaptive PyOD class.

ImportError

If the object appears to require PyOD but PyOD is absent.

TypeError

If required protocol methods are missing.

Source code in nonconform/adapters.py
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
def adapt(detector: Any) -> AnomalyDetector:
    """Return a detector that satisfies the public anomaly-detector protocol.

    Args:
        detector: PyOD, scikit-learn, or custom detector object.

    Returns:
        The original structurally compatible object or a PyOD adapter.

    Raises:
        ValueError: If the detector is a blocked batch-adaptive PyOD class.
        ImportError: If the object appears to require PyOD but PyOD is absent.
        TypeError: If required protocol methods are missing.
    """
    _guard_blocked_pyod_detector(detector)

    if isinstance(detector, AnomalyDetector):
        return detector

    if PYOD_AVAILABLE and isinstance(detector, PyODBaseDetector):
        return PyODAdapter(detector)

    if not PYOD_AVAILABLE and _looks_like_pyod(detector):
        raise ImportError(
            "Detector appears to be a PyOD detector, but PyOD is not installed. "
            'Install with: pip install "nonconform[pyod]" or pip install pyod.'
        )

    required_methods = ["fit", "decision_function", "get_params", "set_params"]
    missing_methods = [m for m in required_methods if not hasattr(detector, m)]
    if missing_methods:
        raise TypeError(
            "Detector must implement AnomalyDetector protocol. "
            f"Missing methods: {', '.join(missing_methods)}"
        )

    return detector

parse_score_polarity

parse_score_polarity(
    score_polarity: ScorePolarityInput,
) -> ScorePolarity

Normalize a score-polarity string or enum.

Parameters:

Name Type Description Default
score_polarity ScorePolarityInput

A :class:ScorePolarity value or one of "auto", "higher_is_anomalous", and "higher_is_normal".

required

Returns:

Type Description
ScorePolarity

The corresponding :class:ScorePolarity member.

Raises:

Type Description
ValueError

If a string value is unsupported.

TypeError

If the input is neither a string nor ScorePolarity.

Source code in nonconform/adapters.py
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
def parse_score_polarity(score_polarity: ScorePolarityInput) -> ScorePolarity:
    """Normalize a score-polarity string or enum.

    Args:
        score_polarity: A :class:`ScorePolarity` value or one of ``"auto"``,
            ``"higher_is_anomalous"``, and ``"higher_is_normal"``.

    Returns:
        The corresponding :class:`ScorePolarity` member.

    Raises:
        ValueError: If a string value is unsupported.
        TypeError: If the input is neither a string nor ``ScorePolarity``.
    """
    if isinstance(score_polarity, ScorePolarity):
        return score_polarity

    if isinstance(score_polarity, str):
        normalized = score_polarity.strip().lower()
        mapping = {
            "auto": ScorePolarity.AUTO,
            "higher_is_anomalous": ScorePolarity.HIGHER_IS_ANOMALOUS,
            "higher_is_normal": ScorePolarity.HIGHER_IS_NORMAL,
        }
        if normalized in mapping:
            return mapping[normalized]
        raise ValueError(
            "Invalid score_polarity value. "
            "Use one of: 'auto', 'higher_is_anomalous', 'higher_is_normal'."
        )

    raise TypeError(
        "score_polarity must be a ScorePolarity enum or string literal "
        "('auto', 'higher_is_anomalous', 'higher_is_normal')."
    )

resolve_implicit_score_polarity

resolve_implicit_score_polarity(
    detector: Any,
) -> ScorePolarity

Resolve score polarity when users omit score_polarity.

The default favors low-friction custom detector onboarding while preserving safe behavior for known detector families: - Known sklearn normality detectors -> HIGHER_IS_NORMAL - PyOD detectors -> HIGHER_IS_ANOMALOUS - Unknown custom detectors -> HIGHER_IS_ANOMALOUS

Parameters:

Name Type Description Default
detector Any

Adapted detector whose family determines the default.

required

Returns:

Type Description
ScorePolarity

The anomalous-higher or normal-higher convention selected by the v1

ScorePolarity

implicit-default policy.

Source code in nonconform/adapters.py
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
def resolve_implicit_score_polarity(detector: Any) -> ScorePolarity:
    """Resolve score polarity when users omit score_polarity.

    The default favors low-friction custom detector onboarding while
    preserving safe behavior for known detector families:
    - Known sklearn normality detectors -> HIGHER_IS_NORMAL
    - PyOD detectors -> HIGHER_IS_ANOMALOUS
    - Unknown custom detectors -> HIGHER_IS_ANOMALOUS

    Args:
        detector: Adapted detector whose family determines the default.

    Returns:
        The anomalous-higher or normal-higher convention selected by the v1
        implicit-default policy.
    """
    if _is_known_sklearn_normality_detector(detector):
        return ScorePolarity.HIGHER_IS_NORMAL
    if isinstance(detector, PyODAdapter) or _looks_like_pyod(detector):
        return ScorePolarity.HIGHER_IS_ANOMALOUS
    return ScorePolarity.HIGHER_IS_ANOMALOUS

resolve_score_polarity

resolve_score_polarity(
    detector: Any, score_polarity: ScorePolarityInput
) -> ScorePolarity

Resolve requested score polarity in strict AUTO mode.

Unlike resolve_implicit_score_polarity, this function is intentionally strict for explicit score_polarity="auto" and raises for unknown detectors.

Parameters:

Name Type Description Default
detector Any

Adapted detector whose family may be inspected.

required
score_polarity ScorePolarityInput

Explicit score convention or "auto".

required

Returns:

Type Description
ScorePolarity

A non-auto score-polarity convention.

Raises:

Type Description
ValueError

If strict automatic inference cannot identify the detector family, or if the value is unsupported.

TypeError

If score_polarity has an unsupported type.

Source code in nonconform/adapters.py
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
def resolve_score_polarity(
    detector: Any,
    score_polarity: ScorePolarityInput,
) -> ScorePolarity:
    """Resolve requested score polarity in strict AUTO mode.

    Unlike ``resolve_implicit_score_polarity``, this function is intentionally
    strict for explicit ``score_polarity="auto"`` and raises for unknown
    detectors.

    Args:
        detector: Adapted detector whose family may be inspected.
        score_polarity: Explicit score convention or ``"auto"``.

    Returns:
        A non-auto score-polarity convention.

    Raises:
        ValueError: If strict automatic inference cannot identify the detector
            family, or if the value is unsupported.
        TypeError: If ``score_polarity`` has an unsupported type.
    """
    parsed = parse_score_polarity(score_polarity)
    if parsed is not ScorePolarity.AUTO:
        return parsed

    if isinstance(detector, PyODAdapter) or _looks_like_pyod(detector):
        return ScorePolarity.HIGHER_IS_ANOMALOUS
    if _is_known_sklearn_normality_detector(detector):
        return ScorePolarity.HIGHER_IS_NORMAL

    detector_cls = type(detector)
    detector_name = f"{detector_cls.__module__}.{detector_cls.__qualname__}"
    raise ValueError(
        "Unable to infer score polarity automatically in strict auto mode for "
        f"detector '{detector_name}'. Auto inference currently supports PyOD "
        "detectors and known sklearn normality estimators. For custom detectors, "
        "pass score_polarity='higher_is_anomalous' (recommended when larger "
        "scores mean more anomalous) or score_polarity='higher_is_normal'."
    )

apply_score_polarity

apply_score_polarity(
    detector: AnomalyDetector,
    score_polarity: ScorePolarityInput,
) -> AnomalyDetector

Normalize detector output so larger values mean more anomalous.

Parameters:

Name Type Description Default
detector AnomalyDetector

Detector whose raw score convention is known.

required
score_polarity ScorePolarityInput

Convention of the detector's raw scores. AUTO must be resolved before this call.

required

Returns:

Type Description
AnomalyDetector

The original detector for anomalous-higher scores, or a polarity adapter

AnomalyDetector

that negates normal-higher scores.

Raises:

Type Description
ValueError

If unresolved automatic polarity is supplied.

Source code in nonconform/adapters.py
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
def apply_score_polarity(
    detector: AnomalyDetector,
    score_polarity: ScorePolarityInput,
) -> AnomalyDetector:
    """Normalize detector output so larger values mean more anomalous.

    Args:
        detector: Detector whose raw score convention is known.
        score_polarity: Convention of the detector's raw scores. ``AUTO`` must
            be resolved before this call.

    Returns:
        The original detector for anomalous-higher scores, or a polarity adapter
        that negates normal-higher scores.

    Raises:
        ValueError: If unresolved automatic polarity is supplied.
    """
    parsed = parse_score_polarity(score_polarity)
    if parsed is ScorePolarity.AUTO:
        raise ValueError(
            "score_polarity='auto' must be resolved first with resolve_score_polarity."
        )
    if parsed is ScorePolarity.HIGHER_IS_ANOMALOUS:
        return detector
    return ScorePolarityAdapter(detector=detector, score_polarity=parsed)