API Reference¶
Reference documentation for the complete v1 public module surface. For the statistical assumptions and decision context behind an API, follow the linked user-guide page before relying on a guarantee.
Start Here¶
If you are looking for task-oriented call sequences, start with Common Workflows.
For the v1 public compatibility contract, see API Stability.
Detector¶
nonconform.detector ¶
Core conformal anomaly detector implementation.
This module provides :class:ConformalDetector, which calibrates scores from a
supported anomaly detector. The default empirical estimator returns rank-based
conformal p-values; other estimators document their own interpretation. Weighted
mode estimates density-ratio weights for covariate-shift workflows. Selection
methods apply a multiple-testing procedure to a complete test batch. The
resulting validity and false discovery rate guarantees depend on the assumptions
documented for the chosen calibration and selection procedure.
Classes:
| Name | Description |
|---|---|
BaseConformalDetector |
Abstract base class for conformal detectors. |
ConformalDetector |
Main conformal anomaly detector with optional weighting. |
BaseConformalDetector ¶
Bases: ABC
Abstract base class for all conformal anomaly detectors.
Defines the core interface that all conformal anomaly detection implementations must provide. Conformal detectors support either an integrated or detached calibration workflow:
- Integrated calibration:
fit()trains detector(s) and computes calibration scores - Detached calibration: train detector externally, then call
calibrate()on a separate calibration dataset - Inference phase:
compute_p_values()applies the configured estimation strategy, whileselect()combines estimation with the configured batch selection procedure
Subclasses must implement both abstract methods.
Note
This is an abstract class and cannot be instantiated directly.
Use ConformalDetector for the main implementation.
fit
abstractmethod
¶
fit(
x: DataFrame | ndarray,
y: ndarray | None = None,
*,
n_jobs: int | None = None,
) -> Self
Fit the detector model(s) and compute calibration scores.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
DataFrame | ndarray
|
The dataset used for fitting the model(s) and determining calibration scores. |
required |
y
|
ndarray | None
|
Ignored. Present for sklearn API compatibility. |
None
|
n_jobs
|
int | None
|
Optional strategy-specific parallelism hint.
Currently used by strategies that expose an |
None
|
Returns:
| Type | Description |
|---|---|
Self
|
The fitted detector instance. |
Source code in nonconform/detector.py
120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 | |
calibrate ¶
calibrate(
x: DataFrame | ndarray, y: ndarray | None = None
) -> Self
Calibrate a pre-fitted detector on separate calibration data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
DataFrame | ndarray
|
Dataset used only to compute calibration scores. |
required |
y
|
ndarray | None
|
Ignored. Present for sklearn API compatibility. |
None
|
Returns:
| Type | Description |
|---|---|
Self
|
The calibrated detector instance. |
Source code in nonconform/detector.py
144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 | |
compute_p_values
abstractmethod
¶
compute_p_values(
x: DataFrame | Series | ndarray,
*,
refit_weights: bool = True,
) -> np.ndarray | pd.Series
Return conformal p-values for new data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
DataFrame | Series | ndarray
|
New data instances for anomaly estimation. |
required |
refit_weights
|
bool
|
Whether to refit the weight estimator for this batch in weighted mode. Ignored in standard mode. |
True
|
Returns:
| Type | Description |
|---|---|
ndarray | Series
|
P-values as ndarray for numpy input, or pandas Series for pandas input. |
Source code in nonconform/detector.py
161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 | |
score_samples
abstractmethod
¶
score_samples(
x: DataFrame | Series | ndarray,
*,
refit_weights: bool = True,
) -> np.ndarray | pd.Series
Return aggregated raw anomaly scores for new data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
DataFrame | Series | ndarray
|
New data instances for anomaly estimation. |
required |
refit_weights
|
bool
|
Whether to refit the weight estimator for this batch in weighted mode. Ignored in standard mode. |
True
|
Returns:
| Type | Description |
|---|---|
ndarray | Series
|
Raw scores as ndarray for numpy input, or pandas Series for pandas input. |
Source code in nonconform/detector.py
180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 | |
ConformalDetector ¶
ConformalDetector(
detector: Any,
strategy: BaseStrategy,
estimation: BaseEstimation | None = None,
weight_estimator: BaseWeightEstimator | None = None,
aggregation: str = "median",
score_polarity: ScorePolarity
| Literal[
"auto", "higher_is_anomalous", "higher_is_normal"
]
| None = None,
seed: int | None = None,
verbose: bool = False,
verify_prepared_batch_content: bool = True,
)
Bases: BaseConformalDetector
Wrap an anomaly detector with conformal calibration and batch selection.
The wrapped detector may be a recognized scikit-learn estimator, a PyOD
model, or a custom object implementing the
:class:~nonconform.structures.AnomalyDetector protocol.
In standard mode, the fitted strategy supplies one or more fixed scoring
rules and calibration-score sets. With the default Empirical estimator,
compute_p_values() ranks each test score against those calibration
scores. The usual marginal p-value validity statement requires
exchangeability of the relevant null calibration and test examples.
select() then applies Benjamini-Hochberg to the full test family. Any FDR
guarantee additionally depends on the dependence assumptions of that
multiple-testing procedure. Other estimation strategies document their own
interpretation and assumptions.
Supplying weight_estimator enables weighted mode. The estimator learns
density-ratio weights from calibration and test covariates, and select()
uses weighted conformalized selection. Its validity requires the applicable
covariate-shift assumptions, support overlap, and adequate weights; the class
cannot verify those scientific assumptions from data alone.
With DerandomizedSplits, fit() retains repeated model/calibration
pairs and select() uniformly aggregates per-split e-values before e-BH.
Inspect last_selection_result for evidence and selection diagnostics.
P-value methods and detached calibration are unavailable for this strategy.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detector
|
Any
|
Anomaly detector (PyOD, sklearn-compatible, or custom). |
required |
strategy
|
BaseStrategy
|
The conformal strategy for fitting, calibration, and evidence construction. DerandomizedSplits selects through e-values and e-BH. |
required |
estimation
|
BaseEstimation | None
|
P-value estimation strategy. Defaults to Empirical(). Unused by DerandomizedSplits, which accepts only None or ordinary Empirical. |
None
|
weight_estimator
|
BaseWeightEstimator | None
|
Weight estimator for covariate shift. Defaults to None. |
None
|
aggregation
|
str
|
Method for aggregating scores from multiple fitted models:
|
'median'
|
score_polarity
|
ScorePolarity | Literal['auto', 'higher_is_anomalous', 'higher_is_normal'] | None
|
Score direction convention. Use |
None
|
seed
|
int | None
|
Random seed for reproducibility. Defaults to None. |
None
|
verbose
|
bool
|
If True, displays aggregation progress for multi-model strategies. Defaults to False. |
False
|
verify_prepared_batch_content
|
bool
|
If True (default), weighted reuse mode
( |
True
|
Attributes:
| Name | Type | Description |
|---|---|---|
detector |
The underlying anomaly detection model. |
|
strategy |
The calibration and evidence-construction strategy. |
|
weight_estimator |
Optional weight estimator for handling covariate shift. |
|
aggregation |
Method for combining scores from multiple models. |
|
score_polarity |
ScorePolarity
|
Resolved score polarity used internally. |
seed |
ScorePolarity
|
Random seed for reproducible results. |
verbose |
ScorePolarity
|
Whether to display progress bars. |
Examples:
Standard conformal p-values and batch selection:
import numpy as np
from sklearn.ensemble import IsolationForest
from nonconform import ConformalDetector, Split
rng = np.random.default_rng(42)
x_reference = rng.normal(size=(300, 2))
x_test = np.vstack([rng.normal(size=(38, 2)), rng.normal(loc=5.0, size=(2, 2))])
detector = ConformalDetector(
detector=IsolationForest(random_state=42),
strategy=Split(n_calib=0.25),
score_polarity="higher_is_normal",
seed=42,
)
detector.fit(x_reference)
p_values = detector.compute_p_values(x_test)
selected = detector.select(x_test, alpha=0.10)
print(p_values.shape, np.flatnonzero(selected))
Weighted conformal p-values under a simulated covariate shift:
import numpy as np
from sklearn.ensemble import IsolationForest
from nonconform import (
ConformalDetector,
Split,
logistic_weight_estimator,
)
rng = np.random.default_rng(7)
x_reference = rng.normal(size=(400, 2))
x_test = rng.normal(loc=0.5, size=(40, 2))
detector = ConformalDetector(
detector=IsolationForest(random_state=7),
strategy=Split(n_calib=0.25),
weight_estimator=logistic_weight_estimator(),
score_polarity="higher_is_normal",
seed=7,
)
detector.fit(x_reference)
p_values = detector.compute_p_values(x_test)
print(p_values.shape, p_values.min(), p_values.max())
Detached calibration with a pre-trained model (Split strategy):
import numpy as np
from sklearn.ensemble import IsolationForest
from nonconform import ConformalDetector, Split
rng = np.random.default_rng(11)
x_fit = rng.normal(size=(200, 2))
x_calibration = rng.normal(size=(100, 2))
x_test = rng.normal(size=(10, 2))
base_detector = IsolationForest(random_state=11).fit(x_fit)
detector = ConformalDetector(
detector=base_detector,
strategy=Split(n_calib=0.2),
score_polarity="higher_is_normal",
)
detector.calibrate(x_calibration)
p_values = detector.compute_p_values(x_test)
print(p_values)
Note
Strict inductive conformal workflows require a fixed training-only score map at inference time. PyOD detectors known to violate this are: CD, COF, COPOD, ECOD, LMDD, LOCI, RGraph, SOD, SOS.
Source code in nonconform/detector.py
348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 | |
detector_set
property
¶
detector_set: list[AnomalyDetector]
Returns a copy of the list of trained detector models.
calibration_set
property
¶
calibration_set: ndarray
Return a copy of calibration scores.
DerandomizedSplits returns a matrix shaped (n_repetitions, n_calibration), with one row per retained model. Other strategies return their existing one-dimensional score array.
calibration_samples
property
¶
calibration_samples: ndarray
Returns a copy of the calibration samples (weighted mode only).
last_result
property
¶
last_result: ConformalResult | None
Return the most recent raw-score or p-value snapshot.
DerandomizedSplits selection instead populates last_selection_result and clears this snapshot. Its raw-score snapshots have no pooled calibration scores.
last_selection_result
property
¶
last_selection_result: EValueSelectionResult | None
Return a defensive snapshot of the latest e-value selection.
None before selection, after fitting or raw scoring, and for existing p-value selection workflows. Arrays remain read-only in the snapshot.
score_polarity
property
¶
score_polarity: ScorePolarity
Returns the resolved score polarity convention.
get_params ¶
get_params(deep: bool = True) -> dict[str, Any]
Return estimator parameters following sklearn conventions.
Notes
deep=Falsereturns constructor-facing parameters used for sklearn clone compatibility.deep=Truealso includes nestedcomponent__paramentries read from the current runtime components (effective/internal state), which may differ from originally passed constructor objects after adaptation/normalization.
Source code in nonconform/detector.py
468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 | |
set_params ¶
set_params(**params: Any) -> Self
Set estimator parameters following sklearn conventions.
Source code in nonconform/detector.py
505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 | |
fit ¶
fit(
x: DataFrame | ndarray,
y: ndarray | None = None,
*,
n_jobs: int | None = None,
) -> Self
Fit detector model(s) and compute calibration scores.
Uses the specified strategy to train the base detector(s) and calculate non-conformity scores on the calibration set.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
DataFrame | ndarray
|
The dataset used for fitting and calibration. |
required |
y
|
ndarray | None
|
Ignored. Present for sklearn API compatibility. |
None
|
n_jobs
|
int | None
|
Optional strategy-specific parallelism hint. Supported by
strategies whose |
None
|
Returns:
| Type | Description |
|---|---|
Self
|
The fitted detector instance (for method chaining). |
Source code in nonconform/detector.py
567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 | |
calibrate ¶
calibrate(
x: DataFrame | ndarray, y: ndarray | None = None
) -> Self
Calibrate a pre-fitted detector on separate calibration data.
This detached workflow is currently supported only for Split strategy,
where a single pre-fitted model is calibrated on a dedicated dataset.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
DataFrame | ndarray
|
Calibration dataset used to compute calibration scores. |
required |
y
|
ndarray | None
|
Ignored. Present for sklearn API compatibility. |
None
|
Returns:
| Type | Description |
|---|---|
Self
|
The calibrated detector instance (for method chaining). |
Raises:
| Type | Description |
|---|---|
ValueError
|
If strategy is not |
NotFittedError
|
If the base detector appears unfitted. |
Source code in nonconform/detector.py
630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 | |
select ¶
select(
x: DataFrame | Series | ndarray,
*,
alpha: float = 0.05,
pruning: Pruning = Pruning.DETERMINISTIC,
seed: int | None = None,
refit_weights: bool = True,
) -> np.ndarray | pd.Series
Construct evidence and select anomalies from one fixed test batch.
This is the single-call batch workflow. It combines
compute_p_values() with Benjamini-Hochberg in standard mode or
weighted conformalized selection in weighted mode. Validity still
depends on the assumptions of both the p-value construction and the
selected multiple-testing procedure.
With DerandomizedSplits, this instead constructs per-split e-values, averages them uniformly, and applies e-BH once. Configure alpha_bh and tie_seed on that strategy. Automatic ties use a separate fitting-derived seed. Selection populates last_selection_result and clears last_result.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
DataFrame | Series | ndarray
|
New data instances for anomaly estimation. |
required |
alpha
|
float
|
Nominal FDR target in |
0.05
|
pruning
|
Pruning
|
Pruning strategy for weighted FDR control. Ignored in
standard (unweighted) mode. Defaults to
|
DETERMINISTIC
|
seed
|
int | None
|
Optional random seed for weighted randomized pruning modes.
When |
None
|
refit_weights
|
bool
|
Whether to refit the weight estimator for this batch in weighted mode. Ignored in standard mode. Defaults to True. |
True
|
Returns:
| Type | Description |
|---|---|
ndarray | Series
|
Boolean selection mask of shape |
ndarray | Series
|
the selected anomaly discoveries. Returns a pandas Series when the |
ndarray | Series
|
input is a DataFrame or Series. |
Examples:
Standard workflow (no weight estimator):
import numpy as np
from sklearn.ensemble import IsolationForest
from nonconform import ConformalDetector, Split
rng = np.random.default_rng(42)
x_reference = rng.normal(size=(300, 2))
x_test = np.vstack(
[rng.normal(size=(38, 2)), rng.normal(loc=5.0, size=(2, 2))]
)
detector = ConformalDetector(
detector=IsolationForest(random_state=42),
strategy=Split(n_calib=0.25),
score_polarity="higher_is_normal",
seed=42,
).fit(x_reference)
selected = detector.select(x_test, alpha=0.10)
print("Selected indices:", np.flatnonzero(selected))
Weighted workflow:
import numpy as np
from sklearn.ensemble import IsolationForest
from nonconform import (
ConformalDetector,
Split,
logistic_weight_estimator,
)
rng = np.random.default_rng(7)
x_reference = rng.normal(size=(400, 2))
x_test = rng.normal(loc=0.5, size=(40, 2))
detector = ConformalDetector(
detector=IsolationForest(random_state=7),
strategy=Split(n_calib=0.25),
weight_estimator=logistic_weight_estimator(),
score_polarity="higher_is_normal",
seed=7,
).fit(x_reference)
selected = detector.select(x_test, alpha=0.10)
print("Number selected:", int(selected.sum()))
Source code in nonconform/detector.py
814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 | |
prepare_weights_for ¶
prepare_weights_for(x: DataFrame | ndarray) -> Self
Prepare weighted conformal state for a specific test batch.
In weighted mode, this fits the weight estimator for the supplied batch without producing predictions. Use this for explicit state transitions in exploratory workflows.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
DataFrame | ndarray
|
Test batch for which weights should be prepared. |
required |
Returns:
| Type | Description |
|---|---|
Self
|
The fitted detector instance (for method chaining). |
Raises:
| Type | Description |
|---|---|
NotFittedError
|
If fit() has not been called. |
RuntimeError
|
If weighted mode is disabled. |
Source code in nonconform/detector.py
946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 | |
score_samples ¶
score_samples(
x: DataFrame | Series | ndarray,
*,
refit_weights: bool = True,
) -> np.ndarray | pd.Series
Return aggregated raw anomaly scores for new data.
Clears last_selection_result. With DerandomizedSplits, raw aggregation is diagnostic only and the last_result snapshot has calib_scores=None; select() instead uses each model's separate score/calibration pair.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
DataFrame | Series | ndarray
|
New data instances for anomaly estimation. |
required |
refit_weights
|
bool
|
Whether to refit the weight estimator for this batch in weighted mode. Defaults to True. |
True
|
Returns:
| Type | Description |
|---|---|
ndarray | Series
|
Aggregated raw anomaly scores. |
Source code in nonconform/detector.py
974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 | |
compute_p_value ¶
compute_p_value(x: Series | ndarray) -> float
Return one value from the configured estimation strategy.
Unavailable for DerandomizedSplits; use select() on a fixed test batch and inspect last_selection_result instead.
This is a single-sample convenience wrapper around
:meth:compute_p_values. It updates :attr:last_result with the
corresponding one-row result and does not update the calibration set.
With the default Empirical estimator, the returned value is a
rank-based conformal p-value. Other estimators define their own
interpretation and assumptions.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
Series | ndarray
|
One-dimensional feature vector with the same number of features used during fitting or detached calibration. |
required |
Returns:
| Type | Description |
|---|---|
float
|
The observation's p-value or score-tail estimate as a Python float. |
Raises:
| Type | Description |
|---|---|
NotFittedError
|
If the detector has not been fitted or calibrated. |
RuntimeError
|
If weighted conformal mode is enabled. Density-ratio estimation requires a representative test batch. |
ValueError
|
If |
Note
Repeated single-sample calls are batch-equivalent only when detector scoring and p-value estimation are sample-wise and deterministic. Randomized tie-breaking can produce different values from one batch call.
Source code in nonconform/detector.py
1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 | |
compute_p_values ¶
compute_p_values(
x: DataFrame | Series | ndarray,
*,
refit_weights: bool = True,
) -> np.ndarray | pd.Series
Return values from the configured estimation strategy for new data.
Unavailable for DerandomizedSplits; use select() and inspect last_selection_result instead.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
DataFrame | Series | ndarray
|
New data instances for anomaly estimation. |
required |
refit_weights
|
bool
|
Whether to refit the weight estimator for this batch in weighted mode. Defaults to True. |
True
|
Returns:
| Type | Description |
|---|---|
ndarray | Series
|
P-values or score-tail estimates. Pandas input produces a Series |
ndarray | Series
|
named |
Source code in nonconform/detector.py
1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 | |
fdp_bounds ¶
fdp_bounds(
x: DataFrame | Series | ndarray,
*,
confidence: float = 0.95,
method: str = "mc_thc",
n_resamples: int | None = None,
seed: int | None = None,
boost: bool = True,
lower: float | None = None,
upper: float | None = None,
beta: float | None = None,
precision: float | None = None,
) -> FDPCertificate
Compute p-values once and return a simultaneous FDP certificate.
Supports unweighted empirical Split inference, including detached calibration. Choose the method before inspecting its curve and keep the testing family fixed. Confidence is coverage, not an FDR target. Scientific exchangeability remains the caller's responsibility.
Options match :meth:nonconform.fdr.FDPCertificate.from_p_values.
The seed controls certificate Monte Carlo sampling only and does not
inherit the fitting seed. The returned certificate is independent of
subsequent detector operations; its select() returns a NumPy mask.
Source code in nonconform/detector.py
1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 1178 1179 1180 1181 1182 1183 1184 1185 | |
Resampling Strategies¶
nonconform.resampling ¶
Calibration strategies for conformal anomaly detection.
These strategies define how detector replicas and calibration scores are formed.
Split uses disjoint fitting and calibration subsets. CrossValidation
uses out-of-fold scores, and JackknifeBootstrap uses out-of-bag scores. The
latter two are package-specific score-aggregation constructions; their names do
not by themselves transfer coverage theorems for conformal prediction intervals
to anomaly p-values.
DerandomizedSplits retains separate held-out calibration rows and constructs
e-values per split before uniform evidence aggregation and e-BH selection.
Classes:
| Name | Description |
|---|---|
BaseStrategy |
Abstract base class for calibration strategies. |
Split |
Simple train-test split strategy. |
DerandomizedSplits |
Repeated split-conformal e-values with e-BH selection. |
CrossValidation |
K-fold cross-validation strategy (includes Jackknife factory). |
JackknifeBootstrap |
Bootstrap out-of-bag calibration strategy. |
BaseStrategy ¶
BaseStrategy(mode: ConformalModeInput = 'plus')
Bases: ABC
Abstract base class for anomaly detection calibration strategies.
This class provides a common interface for various calibration strategies applied to anomaly detectors. Subclasses must implement the core calibration logic and define how calibration data is identified and used.
Attributes:
| Name | Type | Description |
|---|---|---|
_mode |
ConformalMode
|
Model retention mode controlling calibration/inference behavior. |
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
mode
|
ConformalModeInput
|
Model retention mode ( |
'plus'
|
Source code in nonconform/resampling.py
104 105 106 107 108 109 110 111 112 | |
calibration_ids
abstractmethod
property
¶
calibration_ids: list[int] | None
Indices of data points used for calibration.
fit_calibrate
abstractmethod
¶
fit_calibrate(
x: DataFrame | ndarray,
detector: AnomalyDetector,
seed: int | None = None,
weighted: bool = False,
) -> tuple[list[AnomalyDetector], np.ndarray]
Fits the detector and performs calibration.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
DataFrame | ndarray
|
The input data for fitting and calibration. |
required |
detector
|
AnomalyDetector
|
The anomaly detection model to be fitted and calibrated. |
required |
seed
|
int | None
|
Random seed for reproducibility. Defaults to None. |
None
|
weighted
|
bool
|
Whether to use weighted approach. Defaults to False. |
False
|
Returns:
| Type | Description |
|---|---|
tuple[list[AnomalyDetector], ndarray]
|
Tuple of (list of trained detectors, calibration scores array). |
Source code in nonconform/resampling.py
114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 | |
Split ¶
Split(n_calib: float | int = 0.1)
Bases: BaseStrategy
Split conformal strategy for fast anomaly detection.
Implements the classical split conformal approach by dividing training data into separate fitting and calibration sets.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n_calib
|
float | int
|
Size or proportion of data used for calibration. If float, must be between 0.0 and 1.0 (proportion). If int, the absolute number of samples. Defaults to 0.1. |
0.1
|
Examples:
from nonconform import Split
# Use 20% of data for calibration
strategy = Split(n_calib=0.2)
# Use exactly 1000 samples for calibration
strategy = Split(n_calib=1000)
Source code in nonconform/resampling.py
167 168 169 170 | |
calibration_ids
property
¶
calibration_ids: list[int] | None
Indices of calibration samples (None if weighted=False).
fit_calibrate ¶
fit_calibrate(
x: DataFrame | ndarray,
detector: AnomalyDetector,
weighted: bool = False,
seed: int | None = None,
) -> tuple[list[AnomalyDetector], np.ndarray]
Fits detector and generates calibration scores using a data split.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
DataFrame | ndarray
|
The input data. |
required |
detector
|
AnomalyDetector
|
The detector instance to train. |
required |
weighted
|
bool
|
If True, stores calibration sample indices. Defaults to False. |
False
|
seed
|
int | None
|
Random seed for reproducibility. Defaults to None. |
None
|
Returns:
| Type | Description |
|---|---|
tuple[list[AnomalyDetector], ndarray]
|
Tuple of (list with trained detector, calibration scores array). |
Source code in nonconform/resampling.py
200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 | |
DerandomizedSplits ¶
DerandomizedSplits(
n_repetitions: int = 5,
n_calib: float | int = 0.1,
*,
alpha_bh: float | None = None,
tie_seed: int | None = None,
)
Bases: BaseStrategy
Aggregate conformal e-values across repeated random splits with e-BH.
Each replica retains its own held-out calibration scores. Selection converts each replica's scores to e-values before uniformly averaging the evidence; it never pools calibration scores or aggregates raw scores first.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n_repetitions
|
int
|
Positive number of splits. Defaults to five. |
5
|
n_calib
|
float | int
|
Calibration count or fraction, with the same meaning as Split. |
0.1
|
alpha_bh
|
float | None
|
Fixed inner threshold in (0, 1), or None for alpha / 10 using the target supplied to detector.select(). Choose before inspecting the test evidence. |
None
|
tie_seed
|
int | None
|
Optional non-negative override for randomized score ties. None automatically derives a separate random stream during fit. |
None
|
Examples:
from sklearn.ensemble import IsolationForest
from nonconform import ConformalDetector, DerandomizedSplits
detector = ConformalDetector(
detector=IsolationForest(),
strategy=DerandomizedSplits(n_repetitions=5, n_calib=0.2),
seed=42,
)
# detector.fit(x_reference)
# selected = detector.select(x_test, alpha=0.05)
# evidence = detector.last_selection_result.e_values
Note
Requires unweighted, integrated splits and exchangeable normal reference and null test observations. All repetitions score the same fixed test family. The aggregate null-evidence condition supports one final e-BH application; individual values need not be ordinary e-values. Repetition reduces dependence on a particular split but does not remove randomness.
Source code in nonconform/resampling.py
290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 | |
calibration_ids
property
¶
calibration_ids: None
No pooled calibration sample exists for this unweighted strategy.
get_params ¶
get_params(deep: bool = True) -> dict[str, Any]
Return constructor parameters for inspection and sklearn cloning.
Source code in nonconform/resampling.py
309 310 311 312 313 314 315 316 | |
set_params ¶
set_params(**params: Any) -> Self
Update constructor parameters and clear derived random state.
Source code in nonconform/resampling.py
318 319 320 321 322 323 324 325 326 327 328 | |
fit_calibrate ¶
fit_calibrate(
x: DataFrame | ndarray,
detector: AnomalyDetector,
seed: int | None = None,
weighted: bool = False,
) -> tuple[list[AnomalyDetector], np.ndarray]
Fit independent replicas and return aligned calibration-score rows.
Returns:
| Type | Description |
|---|---|
list[AnomalyDetector]
|
Models and scores shaped (n_repetitions, n_calibration). Row i |
ndarray
|
contains only held-out scores produced by model i. |
Source code in nonconform/resampling.py
330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 | |
CrossValidation ¶
CrossValidation(
k: int | None = 5,
mode: ConformalModeInput = "plus",
shuffle: bool = True,
)
Bases: BaseStrategy
K-fold out-of-fold calibration for conformal anomaly scoring.
The strategy trains one detector per fold and records scores for observations
while they are held out. In "plus" mode, test scores are aggregated over
the retained fold models before comparison with the out-of-fold calibration
scores. This is the package's anomaly-score construction, not a claim that a
CV+ prediction-interval theorem applies unchanged.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
k
|
int | None
|
Number of folds. If None, uses leave-one-out (k=n at fit time). |
5
|
mode
|
ConformalModeInput
|
Model retention mode ( |
'plus'
|
shuffle
|
bool
|
Whether to shuffle data before splitting. Defaults to True. Set to False for deterministic leave-one-out (Jackknife). |
True
|
Examples:
from nonconform import CrossValidation
# 5-fold cross-validation
strategy = CrossValidation(k=5)
# Leave-one-out (Jackknife) via factory
strategy = CrossValidation.jackknife()
Source code in nonconform/resampling.py
436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 | |
jackknife
classmethod
¶
jackknife(
mode: ConformalModeInput = "plus",
) -> CrossValidation
Create Leave-One-Out cross-validation (deterministic, no shuffle).
This factory method creates a Jackknife strategy, which is a special case of k-fold CV where k equals n (the dataset size). Each sample is left out exactly once for calibration.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
mode
|
ConformalModeInput
|
Model retention mode ( |
'plus'
|
Returns:
| Type | Description |
|---|---|
CrossValidation
|
CrossValidation configured for leave-one-out. |
Examples:
from nonconform import CrossValidation
strategy = CrossValidation.jackknife()
print(strategy.k, strategy.mode)
Source code in nonconform/resampling.py
462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 | |
fit_calibrate ¶
fit_calibrate(
x: DataFrame | ndarray,
detector: AnomalyDetector,
seed: int | None = None,
weighted: bool = False,
) -> tuple[list[AnomalyDetector], np.ndarray]
Fit and calibrate using k-fold cross-validation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
DataFrame | ndarray
|
Input data matrix. |
required |
detector
|
AnomalyDetector
|
The base anomaly detector. |
required |
seed
|
int | None
|
Random seed for reproducibility. Defaults to None. |
None
|
weighted
|
bool
|
Whether to use weighted calibration. Defaults to False. |
False
|
Returns:
| Type | Description |
|---|---|
tuple[list[AnomalyDetector], ndarray]
|
Tuple of (list of trained detectors, calibration scores array). |
Raises:
| Type | Description |
|---|---|
ValueError
|
If k < 2 or not enough samples for specified k. |
Source code in nonconform/resampling.py
486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 | |
JackknifeBootstrap ¶
JackknifeBootstrap(
n_bootstraps: int = 100,
aggregation_method: BootstrapAggregationMethod = "mean",
mode: ConformalModeInput = "plus",
)
Bases: BaseStrategy
Bootstrap and out-of-bag calibration for conformal anomaly scoring.
Each bootstrap replica is fitted on a sample drawn with replacement. Every
reference observation receives a calibration score aggregated over replicas
for which that observation was out of bag. In "plus" mode, test scores
are aggregated over all retained replicas.
The construction is inspired by jackknife+-after-bootstrap (JaB+), but this class produces anomaly scores and p-values rather than the predictive intervals studied by the JaB+ theorem. Do not infer an interval-coverage guarantee solely from the class name.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n_bootstraps
|
int
|
Number of bootstrap iterations. Defaults to 100. |
100
|
aggregation_method
|
BootstrapAggregationMethod
|
How to aggregate OOB predictions ("mean" or "median"). Defaults to "mean". |
'mean'
|
mode
|
ConformalModeInput
|
Model retention mode ( |
'plus'
|
References
Kim, Byol, Chen Xu, and Rina Foygel Barber. "Predictive Inference Is Free with the Jackknife+-after-Bootstrap." NeurIPS 2020.
Source code in nonconform/resampling.py
643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 | |
calibration_ids
property
¶
calibration_ids: list[int]
Indices used for calibration (all samples in JaB+).
aggregation_method
property
¶
aggregation_method: BootstrapAggregationMethod
Aggregation method for OOB predictions.
fit_calibrate ¶
fit_calibrate(
x: DataFrame | ndarray,
detector: AnomalyDetector,
seed: int | None = None,
weighted: bool = False,
n_jobs: int | None = None,
) -> tuple[list[AnomalyDetector], np.ndarray]
Fit bootstrap replicas and compute out-of-bag calibration scores.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
DataFrame | ndarray
|
Input data matrix. |
required |
detector
|
AnomalyDetector
|
The base anomaly detector. |
required |
seed
|
int | None
|
Random seed for reproducibility. Defaults to None. |
None
|
weighted
|
bool
|
Accepted for the shared strategy interface. Calibration indices already include every input row in this strategy. |
False
|
n_jobs
|
int | None
|
Number of parallel jobs. Use -1 for all available cores. Defaults to None (sequential). |
None
|
Returns:
| Type | Description |
|---|---|
tuple[list[AnomalyDetector], ndarray]
|
Tuple of (list of trained detectors, calibration scores array). |
Source code in nonconform/resampling.py
681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 | |
P-Value Estimation¶
nonconform.scoring ¶
Tail-probability estimation strategies for calibrated anomaly scores.
Empirical implements rank-based conformal p-values. ConditionalEmpirical
adds a calibration map for a stronger conditional target. Probabilistic
instead estimates a smooth score-tail probability with a kernel density model;
it does not inherit the exact finite-sample guarantee of the empirical rank.
Classes:
| Name | Description |
|---|---|
BaseEstimation |
Abstract base class for p-value estimation. |
Empirical |
Classical empirical p-value estimation using discrete CDF. |
ConditionalEmpirical |
Conditionally calibrated empirical p-values. |
Probabilistic |
KDE-based score-tail probability estimation. |
Kernel ¶
Bases: Enum
Kernel functions for KDE-based score-tail estimation.
Attributes:
| Name | Type | Description |
|---|---|---|
GAUSSIAN |
Gaussian (normal) kernel. |
|
EXPONENTIAL |
Exponential kernel. |
|
BOX |
Box (uniform) kernel. |
|
TRIANGULAR |
Triangular kernel. |
|
EPANECHNIKOV |
Epanechnikov kernel. |
|
BIWEIGHT |
Biweight (quartic) kernel. |
|
TRIWEIGHT |
Triweight kernel. |
|
TRICUBE |
Tricube kernel. |
|
COSINE |
Cosine kernel. |
BaseEstimation ¶
Bases: ABC
Abstract base for p-value and score-tail estimation strategies.
compute_p_values
abstractmethod
¶
compute_p_values(
scores: ndarray,
calibration_set: ndarray,
weights: tuple[ndarray, ndarray] | None = None,
) -> np.ndarray
Compute p-values or strategy-defined tail estimates for test scores.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
scores
|
ndarray
|
Test instance anomaly scores (1D array). |
required |
calibration_set
|
ndarray
|
Calibration anomaly scores (1D array). |
required |
weights
|
tuple[ndarray, ndarray] | None
|
Optional (w_calib, w_test) tuple for weighted conformal. |
None
|
Returns:
| Type | Description |
|---|---|
ndarray
|
Array of values for each test instance. The concrete strategy |
ndarray
|
defines their statistical interpretation. |
Source code in nonconform/scoring.py
66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 | |
get_metadata ¶
get_metadata() -> dict[str, Any]
Optional auxiliary data exposed after compute_p_values.
Source code in nonconform/scoring.py
86 87 88 | |
set_seed ¶
set_seed(seed: int | None) -> None
Set random seed for reproducibility.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
seed
|
int | None
|
Random seed value or None. |
required |
Source code in nonconform/scoring.py
90 91 92 93 94 95 96 97 | |
Empirical ¶
Empirical(tie_break: TieBreakModeInput = 'classical')
Bases: BaseEstimation
Classical rank-based empirical conformal p-values.
Computes p-values using deterministic tie handling by default. Optionally supports randomized smoothing of tied calibration mass and the test point's own mass, eliminating the classical resolution floor (Jin & Candes 2023).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tie_break
|
TieBreakModeInput
|
Tie-breaking strategy ( |
'classical'
|
Examples:
import numpy as np
from nonconform import Empirical
calibration_scores = np.array([0.2, 0.5, 0.7, 1.1, 1.4])
test_scores = np.array([0.6, 1.5])
estimation = Empirical() # tie_break="classical" by default
p_values = estimation.compute_p_values(test_scores, calibration_scores)
print(p_values)
# For randomized smoothing:
randomized = Empirical(tie_break="randomized")
randomized.set_seed(42)
print(randomized.compute_p_values(test_scores, calibration_scores))
Source code in nonconform/scoring.py
130 131 132 | |
set_seed ¶
set_seed(seed: int | None) -> None
Set random seed for reproducibility.
Source code in nonconform/scoring.py
134 135 136 | |
compute_p_values ¶
compute_p_values(
scores: ndarray,
calibration_set: ndarray,
weights: tuple[ndarray, ndarray] | None = None,
) -> np.ndarray
Compute empirical p-values from calibration set.
Source code in nonconform/scoring.py
138 139 140 141 142 143 144 145 146 147 148 149 | |
ConditionalEmpirical ¶
ConditionalEmpirical(
*,
delta: float = 0.05,
method: str | ConditionalCalibrationMethod = "mc",
tie_break: TieBreakModeInput = "classical",
simes_kden: int = 2,
mc_num_simulations: int = 10000,
)
Bases: Empirical
Conditionally calibrated empirical conformal p-values (CCCPV).
This estimator first computes classical empirical conformal p-values and then applies a finite-sample calibration map:
.. math:: p_j = \frac{1 + \sum_{i=1}^{n_{\text{cal}}}\mathbf{1}[s_i \ge s_j]} {n_{\text{cal}} + 1}, \qquad \tilde p_j = C_{n_{\text{cal}},\delta}(p_j).
Supported calibration maps are "mc", "simes", "dkwm", and
"asymptotic".
References
Bates et al. (2023), Testing for outliers with conformal p-values. Reference implementation: https://github.com/msesia/conditional-conformal-pvalues
Note
Weighted conformal p-values are intentionally not supported in this
first release of ConditionalEmpirical.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
delta
|
float
|
Failure-probability parameter used by the conditional calibration
map; the associated calibration event has probability at least
|
0.05
|
method
|
str | ConditionalCalibrationMethod
|
Conditional calibration method. One of
|
'mc'
|
tie_break
|
TieBreakModeInput
|
Tie-breaking strategy used for base empirical p-values
( |
'classical'
|
simes_kden
|
int
|
Denominator used to derive |
2
|
mc_num_simulations
|
int
|
Monte Carlo sample size used to estimate the
finite-sample correction for |
10000
|
Source code in nonconform/scoring.py
221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 | |
set_seed ¶
set_seed(seed: int | None) -> None
Set random seed for reproducibility.
Source code in nonconform/scoring.py
256 257 258 259 260 | |
compute_p_values ¶
compute_p_values(
scores: ndarray,
calibration_set: ndarray,
weights: tuple[ndarray, ndarray] | None = None,
) -> np.ndarray
Compute conditionally calibrated conformal p-values.
Source code in nonconform/scoring.py
262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 | |
Probabilistic ¶
Probabilistic(
kernel: Kernel | Sequence[Kernel] = Kernel.GAUSSIAN,
n_trials: int = 100,
cv_folds: int = -1,
)
Bases: BaseEstimation
KDE-based continuous estimates of score-tail probability.
Fits a kernel density estimate to calibration scores and evaluates its survival function. The results are model-based estimates, not rank-based conformal p-values, and therefore do not have the empirical estimator's exact finite-sample distribution-free guarantee. Smoothness or finer numerical resolution is not evidence of calibration.
The estimator supports automatic hyperparameter tuning and calibration weights. In weighted mode, only calibration weights are applied to the KDE; test weights are intentionally not injected into the survival calculation. This is not the exact discrete weighted conformal construction, and it lets tail estimates reach 0 instead of imposing the lower bound w_test / (sum_calib_weight + w_test) that the discrete weighted formula would impose.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
kernel
|
Kernel | Sequence[Kernel]
|
Kernel function or list (list triggers kernel tuning). Bandwidth is always auto-tuned. Defaults to Kernel.GAUSSIAN. |
GAUSSIAN
|
n_trials
|
int
|
Number of Optuna trials for tuning. Defaults to 100. |
100
|
cv_folds
|
int
|
CV folds for tuning (-1 for leave-one-out). Defaults to -1. |
-1
|
Examples:
import numpy as np
from nonconform import Probabilistic
from nonconform.enums import Kernel
rng = np.random.default_rng(42)
calibration_scores = rng.normal(size=200)
test_scores = np.array([0.0, 1.0, 2.0])
# n_trials=0 skips optional hyperparameter search for a quick example.
estimation = Probabilistic(kernel=Kernel.GAUSSIAN, n_trials=0)
tail_estimates = estimation.compute_p_values(test_scores, calibration_scores)
print(tail_estimates)
Source code in nonconform/scoring.py
489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 | |
compute_p_values ¶
compute_p_values(
scores: ndarray,
calibration_set: ndarray,
weights: tuple[ndarray, ndarray] | None = None,
) -> np.ndarray
Compute continuous p-values using KDE.
Lazy fitting: tunes and fits KDE on first call or when calibration changes. Note: When weights are provided, this estimator uses only calibration weights to shape the KDE. Test weights are accepted for API parity but do not set a positive lower bound on p-values.
Source code in nonconform/scoring.py
507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 | |
get_metadata ¶
get_metadata() -> dict[str, Any]
Return KDE metadata after p-value computation.
Source code in nonconform/scoring.py
628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 | |
calculate_p_val ¶
calculate_p_val(
scores: ndarray,
calibration_set: ndarray,
tie_break: TieBreakModeInput = "classical",
rng: Generator | None = None,
) -> np.ndarray
Calculate empirical p-values (standalone function).
Uses classical deterministic tie handling by default. Randomized mode interpolates tied calibration mass and the test point's own mass, removing the classical positive resolution floor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
scores
|
ndarray
|
Test instance anomaly scores (1D array). |
required |
calibration_set
|
ndarray
|
Calibration anomaly scores (1D array). |
required |
tie_break
|
TieBreakModeInput
|
Tie-breaking strategy for equal scores ( |
'classical'
|
rng
|
Generator | None
|
Optional random number generator for reproducibility. |
None
|
Returns:
| Type | Description |
|---|---|
ndarray
|
Array of p-values for each test instance. |
Source code in nonconform/scoring.py
299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 | |
calculate_weighted_p_val ¶
calculate_weighted_p_val(
scores: ndarray,
calibration_set: ndarray,
test_weights: ndarray,
calib_weights: ndarray,
tie_break: TieBreakModeInput = "classical",
rng: Generator | None = None,
) -> np.ndarray
Calculate weighted empirical p-values (standalone function).
The default "classical" mode uses deterministic discrete tie handling:
calibration mass at or above the test score is included, together with the
test point's own weight. For calibration mass W_cal and test score
s with test weight w_test, this is
(W_{>=}(s) + w_test) / (W_cal + w_test). This variant is conservative
for discrete scores and agrees with the unweighted classical formula when
all weights are one. Without calibration scores tied with s, it equals
the strict deterministic formula used by Jin and Candes (2023).
The "randomized" mode uses the tie-safe weighted conformal formula
(W_{>}(s) + U * (W_{=}(s) + w_test)) / (W_cal + w_test), where
U is uniform on [0, 1]. It avoids the discrete resolution floor and
interpolates both tied calibration mass and the test point's own mass.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
scores
|
ndarray
|
Test instance anomaly scores (1D array). |
required |
calibration_set
|
ndarray
|
Calibration anomaly scores (1D array). |
required |
test_weights
|
ndarray
|
Test instance weights (1D array). |
required |
calib_weights
|
ndarray
|
Calibration weights (1D array). |
required |
tie_break
|
TieBreakModeInput
|
Tie-breaking strategy for equal scores ( |
'classical'
|
rng
|
Generator | None
|
Optional random number generator for reproducibility. |
None
|
Returns:
| Type | Description |
|---|---|
ndarray
|
Array of weighted p-values for each test instance. |
Note
Classical mode has the positive lower bound
test_weights / (sum(calib_weights) + test_weights) when no
calibration mass lies above the test score. Randomized mode has no such
positive floor because it multiplies the test point's mass by U.
Source code in nonconform/scoring.py
344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 | |
Weight Estimation¶
nonconform.weighting ¶
Weight estimation for covariate-shift workflows in weighted conformal inference.
These estimators approximate density ratios
w(x) = p_test(x) / p_calibration(x) from calibration and target covariates.
Weighted conformal methods can use those ratios under a covariate-shift model,
support overlap, and the method's other assumptions. Estimation, misspecification,
and clipping can all affect the resulting validity; this module does not infer a
distribution-shift guarantee merely by fitting a classifier.
Classes:
| Name | Description |
|---|---|
BaseWeightEstimator |
Abstract base class for weight estimators. |
IdentityWeightEstimator |
Returns uniform weights (no covariate shift). |
SklearnWeightEstimator |
Wrapper for sklearn probabilistic classifiers. |
BootstrapBaggedWeightEstimator |
Bootstrap-bagged, batch-specific estimator. |
Factory functions
logistic_weight_estimator: Create estimator using Logistic Regression. forest_weight_estimator: Create estimator using Random Forest.
ProbabilisticClassifier ¶
Bases: Protocol
Protocol for classifiers that support probability estimation.
This protocol defines the interface for sklearn-compatible classifiers that can produce probability estimates for weight computation.
fit ¶
fit(X: ndarray, y: ndarray) -> ProbabilisticClassifier
Fit the classifier on training data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
ndarray
|
Feature matrix of shape (n_samples, n_features). |
required |
y
|
ndarray
|
Target labels of shape (n_samples,). |
required |
Returns:
| Type | Description |
|---|---|
ProbabilisticClassifier
|
The fitted classifier instance. |
Source code in nonconform/weighting.py
51 52 53 54 55 56 57 58 59 60 61 | |
predict_proba ¶
predict_proba(X: ndarray) -> np.ndarray
Return probability estimates for samples.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
ndarray
|
Feature matrix of shape (n_samples, n_features). |
required |
Returns:
| Type | Description |
|---|---|
ndarray
|
Probability estimates of shape (n_samples, n_classes). |
Source code in nonconform/weighting.py
63 64 65 66 67 68 69 70 71 72 | |
BaseWeightEstimator ¶
Bases: ABC
Abstract base class for weighted conformal density-ratio estimators.
Weight estimators approximate the relative density of target covariates with respect to calibration covariates. A downstream statistical guarantee also requires an appropriate shift model, overlap, and adequate ratios.
Subclasses must implement fit(), _get_stored_weights(), and _score_new_data() to provide specific weight estimation strategies.
fit
abstractmethod
¶
fit(
calibration_samples: ndarray, test_samples: ndarray
) -> None
Estimate density ratio weights.
Source code in nonconform/weighting.py
88 89 90 91 | |
get_weights ¶
get_weights(
calibration_samples: ndarray | None = None,
test_samples: ndarray | None = None,
) -> tuple[np.ndarray, np.ndarray]
Return density ratio weights for calibration and test data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
calibration_samples
|
ndarray | None
|
Optional calibration data to score. If provided, computes weights for this data using the fitted model. If None, returns stored weights from fit(). Must provide both or neither. |
None
|
test_samples
|
ndarray | None
|
Optional test data to score. If provided, computes weights for this data using the fitted model. If None, returns stored weights from fit(). Must provide both or neither. |
None
|
Returns:
| Type | Description |
|---|---|
tuple[ndarray, ndarray]
|
Tuple of (calibration_weights, test_weights) as numpy arrays. |
Raises:
| Type | Description |
|---|---|
NotFittedError
|
If fit() has not been called. |
ValueError
|
If only one of calibration_samples/test_samples is provided. |
Source code in nonconform/weighting.py
93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 | |
set_seed ¶
set_seed(seed: int | None) -> None
Set random seed for reproducibility.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
seed
|
int | None
|
Random seed value or None. |
required |
Source code in nonconform/weighting.py
141 142 143 144 145 146 147 | |
IdentityWeightEstimator ¶
IdentityWeightEstimator()
Bases: BaseWeightEstimator
Identity weight estimator that returns uniform weights.
This estimator assumes no covariate shift and returns weights of 1.0 for all samples. It is a baseline that deliberately applies no density-ratio correction.
With Empirical, unit weights reproduce the unweighted p-value formula.
Configuring this object as ConformalDetector.weight_estimator still puts
the detector in weighted mode, so select() uses weighted conformalized
selection rather than ordinary Benjamini-Hochberg.
Source code in nonconform/weighting.py
244 245 246 247 | |
fit ¶
fit(
calibration_samples: ndarray, test_samples: ndarray
) -> None
Fit the identity weight estimator.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
calibration_samples
|
ndarray
|
Array of calibration data samples. |
required |
test_samples
|
ndarray
|
Array of test data samples. |
required |
Source code in nonconform/weighting.py
249 250 251 252 253 254 255 256 257 258 | |
SklearnWeightEstimator ¶
SklearnWeightEstimator(
base_estimator: ProbabilisticClassifier
| BaseEstimator
| None = None,
clip_quantile: float | None = 0.05,
)
Bases: BaseWeightEstimator
Wrap an sklearn-compatible probabilistic binary classifier.
The configured estimator is cloned before fitting. It must expose
fit(), predict_proba(), and classes_ after fitting.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
base_estimator
|
ProbabilisticClassifier | BaseEstimator | None
|
Configured sklearn classifier instance with predict_proba support. Defaults to LogisticRegression. |
None
|
clip_quantile
|
float | None
|
Quantile for weight clipping (e.g., 0.05 clips to 5th-95th percentile). Use None to disable clipping. Defaults to 0.05. |
0.05
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If base_estimator does not implement predict_proba. |
Examples:
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from nonconform.weighting import SklearnWeightEstimator
rng = np.random.default_rng(42)
calibration_samples = rng.normal(size=(120, 3))
test_samples = rng.normal(loc=0.5, size=(80, 3))
estimator = SklearnWeightEstimator(
base_estimator=make_pipeline(
StandardScaler(), LogisticRegression(C=1.0, class_weight="balanced")
)
)
estimator.set_seed(42)
estimator.fit(calibration_samples, test_samples)
calibration_weights, test_weights = estimator.get_weights()
print(calibration_weights.shape, test_weights.shape)
Source code in nonconform/weighting.py
318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 | |
fit ¶
fit(
calibration_samples: ndarray, test_samples: ndarray
) -> None
Fit the weight estimator on calibration and test samples.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
calibration_samples
|
ndarray
|
Array of calibration data samples. |
required |
test_samples
|
ndarray
|
Array of test data samples. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If calibration_samples is empty. |
Source code in nonconform/weighting.py
354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 | |
BootstrapBaggedWeightEstimator ¶
BootstrapBaggedWeightEstimator(
base_estimator: BaseWeightEstimator,
n_bootstraps: int = 100,
clip_quantile: float | None = 0.05,
scoring_mode: Literal["frozen"] = "frozen",
)
Bases: BaseWeightEstimator
Bootstrap-bagged wrapper for weight estimators with instance-wise aggregation.
This estimator repeatedly refits a base weight estimator on balanced bootstrap samples. Geometric averaging can reduce estimator variability in some settings, but it is not universally more accurate or more valid. Compare it against an unbagged estimator on a design representative of deployment.
The algorithm: 1. For each bootstrap iteration: - Resample BOTH sets to balanced sample size (min of calibration and test sizes) - Fit the base estimator on the balanced bootstrap sample - Score ALL original instances using the fitted model (perfect coverage) - Store log(weights) for each instance 2. After all iterations: - Aggregate instance-wise weights using geometric mean (average in log-space) - Optionally clip the aggregated weights at empirical quantiles
Clipping is a numerical stabilization choice that changes the estimated density ratio. Any claimed guarantee must match the clipped construction; a finite observed range alone does not establish the required assumptions.
Seed inheritance
This class uses the _seed attribute pattern for automatic seed
inheritance from ConformalDetector.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
base_estimator
|
BaseWeightEstimator
|
Any BaseWeightEstimator instance. |
required |
n_bootstraps
|
int
|
Number of bootstrap iterations. Defaults to 100. |
100
|
clip_quantile
|
float | None
|
Quantile for adaptive clipping. Use None to disable clipping. Defaults to 0.05. |
0.05
|
scoring_mode
|
Literal['frozen']
|
Weight scoring behavior after fit. Currently only
|
'frozen'
|
Source code in nonconform/weighting.py
473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 | |
supports_rescoring
property
¶
supports_rescoring: bool
Whether this estimator can score arbitrary new batches after fit().
weight_counts
property
¶
weight_counts: str
Return diagnostic info about instance-wise weight coverage.
fit ¶
fit(
calibration_samples: ndarray, test_samples: ndarray
) -> None
Fit the bagged weight estimator with perfect instance coverage.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
calibration_samples
|
ndarray
|
Array of calibration data samples. |
required |
test_samples
|
ndarray
|
Array of test data samples. |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If calibration_samples is empty. |
Source code in nonconform/weighting.py
521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 | |
logistic_weight_estimator ¶
logistic_weight_estimator(
regularization: str | float = "auto",
clip_quantile: float = 0.05,
class_weight: str | dict = "balanced",
max_iter: int = 1000,
) -> SklearnWeightEstimator
Create weight estimator using Logistic Regression.
Note
When used with ConformalDetector, the detector's seed is automatically propagated to the weight estimator for reproducibility.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
regularization
|
str | float
|
Regularization parameter. If 'auto', uses C=1.0. If float, uses as C parameter. |
'auto'
|
clip_quantile
|
float
|
Quantile for weight clipping. Defaults to 0.05. |
0.05
|
class_weight
|
str | dict
|
Class weights for LogisticRegression. Defaults to 'balanced'. |
'balanced'
|
max_iter
|
int
|
Maximum iterations for solver convergence. Defaults to 1000. |
1000
|
Returns:
| Type | Description |
|---|---|
SklearnWeightEstimator
|
Configured SklearnWeightEstimator instance. |
Examples:
import numpy as np
from nonconform import logistic_weight_estimator
rng = np.random.default_rng(42)
calibration_samples = rng.normal(size=(120, 3))
test_samples = rng.normal(loc=0.5, size=(80, 3))
estimator = logistic_weight_estimator(regularization=0.5)
estimator.set_seed(42)
estimator.fit(calibration_samples, test_samples)
calibration_weights, test_weights = estimator.get_weights()
print(calibration_weights.shape, test_weights.shape)
Source code in nonconform/weighting.py
648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 | |
forest_weight_estimator ¶
forest_weight_estimator(
n_estimators: int = 100,
max_depth: int | None = 5,
min_samples_leaf: int = 10,
clip_quantile: float = 0.05,
) -> SklearnWeightEstimator
Create weight estimator using Random Forest.
Note
When used with ConformalDetector, the detector's seed is automatically propagated to the weight estimator for reproducibility.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n_estimators
|
int
|
Number of trees in the forest. Defaults to 100. |
100
|
max_depth
|
int | None
|
Maximum depth of trees. Defaults to 5. |
5
|
min_samples_leaf
|
int
|
Minimum samples at leaf node. Defaults to 10. |
10
|
clip_quantile
|
float
|
Quantile for weight clipping. Defaults to 0.05. |
0.05
|
Returns:
| Type | Description |
|---|---|
SklearnWeightEstimator
|
Configured SklearnWeightEstimator instance. |
Examples:
import numpy as np
from nonconform import forest_weight_estimator
rng = np.random.default_rng(42)
calibration_samples = rng.normal(size=(120, 3))
test_samples = rng.normal(loc=0.5, size=(80, 3))
estimator = forest_weight_estimator(n_estimators=50)
estimator.set_seed(42)
estimator.fit(calibration_samples, test_samples)
calibration_weights, test_weights = estimator.get_weights()
print(calibration_weights.shape, test_weights.shape)
Source code in nonconform/weighting.py
703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 | |
FDR Control¶
Includes post-hoc FDP bounds, derandomized e-value selection, and weighted
low-level expert APIs (weighted_false_discovery_control).
For batch workflows, prefer ConformalDetector.select(...). With
DerandomizedSplits, it applies e-BH and exposes evidence through
last_selection_result; standalone e-value functions remain available for
expert use.
For simultaneous realized-FDP certification, use detector.fdp_bounds(x, ...)
or result.fdp_bounds(...); both return an immutable FDPCertificate.
FDPCertificate.from_p_values(...) is the explicit expert array interface.
See the FDP migration notes
for the intentional replacement of the previous FDP-specific API.
nonconform.fdr ¶
Public false-discovery procedures for conformal anomaly evidence.
Pruning ¶
Bases: Enum
Pruning strategies for weighted FDR control.
Attributes:
| Name | Type | Description |
|---|---|---|
HETEROGENEOUS |
Use independent uniform draws for candidate-specific randomized WCS pruning. |
|
HOMOGENEOUS |
Use one shared uniform draw for randomized WCS pruning. |
|
DETERMINISTIC |
Use the non-randomized WCS pruning rule. |
EValueSelectionResult
dataclass
¶
EValueSelectionResult(
e_values: ndarray,
selected: ndarray,
alpha: float,
alpha_bh: float,
e_threshold: float,
n_repetitions: int,
n_calibration: int,
tie_seed: int | None,
)
Batch e-value FDR selection result.
Attributes:
| Name | Type | Description |
|---|---|---|
e_values |
ndarray
|
Uniformly aggregated conformal e-values. |
selected |
ndarray
|
Boolean e-BH discovery mask. |
alpha |
float
|
Target FDR level supplied to e-BH. |
alpha_bh |
float
|
Fixed inner threshold used to construct split e-values. |
e_threshold |
float
|
Selected e-value cutoff, or infinity when none are selected. |
n_repetitions |
int
|
Number of split-conformal results aggregated. |
n_calibration |
int
|
Number of calibration scores in every repetition. |
tie_seed |
int | None
|
Seed used for randomized score ties, or None when ties were rejected. |
FDPCertificate
dataclass
¶
FDPCertificate()
Immutable simultaneous certificate for realized FDP at p-value cutoffs.
Construct via detector.fdp_bounds(x), result.fdp_bounds(), or the
expert from_p_values() factory. Choose the envelope method before
inspecting its curve. Thresholds may then be explored within this fixed
testing family. Confidence is simultaneous coverage, not an FDR target.
Evidence and default-grid diagnostics are read-only arrays. Queries never
resample. select(t) returns an original-order NumPy mask for p <= t;
t is a p-value cutoff, not a requested FDP bound.
Source code in nonconform/fdr.py
74 75 76 77 78 79 | |
thresholds
property
¶
thresholds: ndarray
Read-only default grid of sorted unique observed p-values.
rejection_counts
property
¶
rejection_counts: ndarray
Read-only discovery counts on the default grid.
fdp_upper_bounds
property
¶
fdp_upper_bounds: ndarray
Read-only FDP upper bounds on the default grid.
precision_lower_bounds
property
¶
precision_lower_bounds: ndarray
Read-only precision lower bounds on the default grid.
n_resamples
property
¶
n_resamples: int | None
Effective Monte Carlo draws; None for deterministic KS.
from_p_values
classmethod
¶
from_p_values(
p_values: ndarray,
*,
n_calibration: int,
confidence: float = 0.95,
method: str = "mc_thc",
n_resamples: int | None = None,
seed: int | None = None,
boost: bool = True,
lower: float | None = None,
upper: float | None = None,
beta: float | None = None,
precision: float | None = None,
) -> FDPCertificate
Certify external p-values; the caller owns provenance assumptions.
Requires unweighted empirical split-conformal p-values from a fixed scoring map and the reference method's exchangeability assumptions. Native detector/snapshot entry points check supported scope; this expert route cannot. Scientific exchangeability is never established by code.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
p_values
|
ndarray
|
Nonempty 1D testing family in [0, 1], in original order. |
required |
n_calibration
|
int
|
Positive calibration size shared by all p-values. |
required |
confidence
|
float
|
Simultaneous coverage probability in (0, 1). |
0.95
|
method
|
str
|
mc_thc (default), mc_hc, mc_ks, ks, or mc_bj. |
'mc_thc'
|
n_resamples
|
int | None
|
Monte Carlo draws; defaults to 1000 for MC methods. |
None
|
seed
|
int | None
|
Monte Carlo seed only. None draws fresh randomness once. |
None
|
boost
|
bool
|
Apply threshold-specific sharpening (default True). |
True
|
lower
|
float | None
|
THC lower truncation, default 0.01. |
None
|
upper
|
float | None
|
THC upper truncation, default 0.99. |
None
|
beta
|
float | None
|
THC exponent, default 0.5. |
None
|
precision
|
float | None
|
BJ inversion tolerance, default 1e-8. |
None
|
Method-specific options must be omitted or None when inapplicable. Deterministic ks accepts neither n_resamples nor seed.
References
Song, Jin, and Candès, "Everywhere Valid Bounds on False Discovery Proportions in Conformal Inference" (2026), arXiv:2605.20726.
Source code in nonconform/fdr.py
97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 | |
bound_at ¶
bound_at(threshold: float | ndarray) -> float | np.ndarray
Evaluate the simultaneous FDP bound at scalar or vector cutoffs.
Source code in nonconform/fdr.py
193 194 195 196 197 | |
precision_at ¶
precision_at(
threshold: float | ndarray,
) -> float | np.ndarray
Return 1 - bound_at(threshold), a simultaneous precision lower bound.
Source code in nonconform/fdr.py
199 200 201 | |
select ¶
select(threshold: float) -> np.ndarray
Return an original-order Boolean NumPy mask for p <= threshold.
Source code in nonconform/fdr.py
203 204 205 206 207 208 | |
to_frame ¶
to_frame(thresholds: ndarray | None = None) -> pd.DataFrame
Report a grid; default to sorted unique observed p-values.
Explicit grids preserve order and duplicates and may be empty. Returned tables are independent of the certificate.
Source code in nonconform/fdr.py
210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 | |
conformal_e_values ¶
conformal_e_values(
test_scores: ndarray,
calib_scores: ndarray,
*,
alpha_bh: float,
tie_seed: int | None = None,
) -> np.ndarray
Compute derandomized conformal e-values from split-conformal scores.
This low-level array interface trusts the caller to provide repetitions for the same test family in the same observation order. Repetitions are aggregated uniformly.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
test_scores
|
ndarray
|
Test anomaly scores. Shape |
required |
calib_scores
|
ndarray
|
Calibration anomaly scores with matching split dimension. |
required |
alpha_bh
|
float
|
Inner BH-style threshold for each split construction. |
required |
tie_seed
|
int | None
|
|
None
|
Returns:
| Type | Description |
|---|---|
ndarray
|
Aggregated e-values of shape |
Raises:
| Type | Description |
|---|---|
TypeError
|
If |
ValueError
|
If score inputs, |
Source code in nonconform/fdr.py
312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 | |
e_value_false_discovery_control ¶
e_value_false_discovery_control(
e_values: ndarray, *, alpha: float = 0.05
) -> np.ndarray
Apply the e-BH procedure to non-negative e-values.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
e_values
|
ndarray
|
Non-negative e-values; larger values are stronger evidence. |
required |
alpha
|
float
|
Target FDR level in |
0.05
|
Returns:
| Type | Description |
|---|---|
ndarray
|
Boolean selection mask aligned with |
Source code in nonconform/fdr.py
349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 | |
select_conformal_e_values ¶
select_conformal_e_values(
results: Sequence[ConformalResult],
*,
alpha: float = 0.05,
alpha_bh: float | None = None,
tie_seed: int | None = None,
) -> EValueSelectionResult
Select a fixed test family from repeated split-conformal results.
Native detector provenance checks integrated, unweighted Split results
and the recorded test-batch content and ordering. Supply unmodified snapshots:
changes to their score arrays are not tracked. Manual or unstamped results
are unsupported; expert callers can use :func:conformal_e_values directly.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
results
|
Sequence[ConformalResult]
|
Non-empty sequence of unmodified detector-produced snapshots. |
required |
alpha
|
float
|
Target FDR level for the final e-BH procedure. |
0.05
|
alpha_bh
|
float | None
|
Inner threshold, defaulting to |
None
|
tie_seed
|
int | None
|
|
None
|
Returns:
| Type | Description |
|---|---|
EValueSelectionResult
|
Aggregated e-values, final e-BH mask, and diagnostics. |
Raises:
| Type | Description |
|---|---|
TypeError
|
If |
ValueError
|
If provenance, batch identity, score arrays, probabilities, or tied-score handling are unsupported. |
Source code in nonconform/fdr.py
369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 | |
weighted_false_discovery_control ¶
weighted_false_discovery_control(
result: ConformalResult | None,
*,
alpha: float = 0.05,
pruning: Pruning = Pruning.DETERMINISTIC,
seed: int | None = None,
) -> np.ndarray
Apply weighted conformalized selection to a result bundle.
The result must contain p-values, test and calibration scores, and matching non-negative weights for the same complete testing family. Validity also depends on the weighted-conformal covariate-shift assumptions.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
result
|
ConformalResult | None
|
Weighted detector result for the target family. |
required |
alpha
|
float
|
Nominal FDR target in |
0.05
|
pruning
|
Pruning
|
Deterministic, homogeneous-randomized, or heterogeneous-randomized WCS pruning rule. |
DETERMINISTIC
|
seed
|
int | None
|
Non-negative seed for randomized pruning, or None. |
None
|
Returns:
| Type | Description |
|---|---|
ndarray
|
Boolean selection mask aligned with the result's test rows. |
Source code in nonconform/fdr.py
442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 | |
weighted_false_discovery_control_from_arrays ¶
weighted_false_discovery_control_from_arrays(
*,
p_values: ndarray,
test_scores: ndarray,
calib_scores: ndarray,
test_weights: ndarray,
calib_weights: ndarray,
alpha: float = 0.05,
pruning: Pruning = Pruning.DETERMINISTIC,
seed: int | None = None,
) -> np.ndarray
Apply weighted conformalized selection to explicit arrays.
This low-level API cannot verify provenance. All arrays must come from the same calibration construction and complete target family.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
p_values
|
ndarray
|
One p-value per test observation. |
required |
test_scores
|
ndarray
|
One anomalous-higher score per test observation. |
required |
calib_scores
|
ndarray
|
Calibration scores in the same orientation. |
required |
test_weights
|
ndarray
|
Non-negative target-density weights for test observations. |
required |
calib_weights
|
ndarray
|
Non-negative target-density weights for calibration observations. |
required |
alpha
|
float
|
Nominal FDR target in |
0.05
|
pruning
|
Pruning
|
WCS pruning rule. |
DETERMINISTIC
|
seed
|
int | None
|
Non-negative seed for randomized pruning, or None. |
None
|
Returns:
| Type | Description |
|---|---|
ndarray
|
Boolean selection mask aligned with the test arrays. |
Source code in nonconform/fdr.py
483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 | |
Martingales¶
nonconform.martingales ¶
Exchangeability martingales for sequential conformal evidence.
This module implements p-value-based martingales and alarm statistics for streaming or temporal monitoring workflows. In practice, you feed one conformal p-value at a time and read a running evidence state after each update.
Implemented martingales
- PowerMartingale
- SimpleMixtureMartingale
- SimpleJumperMartingale
- MixtureMartingale
All classes consume conformal p-values in [0, 1]. Alarm statistics are
computed from martingale ratio increments or mixed component statistics and
exposed together with the current martingale value in :class:MartingaleState.
AlarmConfig
dataclass
¶
AlarmConfig(
ville_threshold: float | None = None,
restarted_ville_threshold: float | None = None,
cusum_threshold: float | None = None,
shiryaev_roberts_threshold: float | None = None,
)
Optional alarm thresholds for martingale evidence statistics.
Thresholds are disabled when set to None. Each threshold compares
against a running statistic in :class:MartingaleState.
ville_threshold and restarted_ville_threshold are Ville thresholds
for e-processes. cusum_threshold and shiryaev_roberts_threshold are
change-evidence thresholds and should not be interpreted as
probability-of-ever-crossing Ville thresholds without a separate theorem for
the exact statistic.
Attributes:
| Name | Type | Description |
|---|---|---|
ville_threshold |
float | None
|
Threshold for the original cumulative martingale. |
restarted_ville_threshold |
float | None
|
Threshold for the harmonic restart-mixture e-process. |
cusum_threshold |
float | None
|
Threshold for the multiplicative CUSUM statistic. |
shiryaev_roberts_threshold |
float | None
|
Threshold for the Shiryaev-Roberts statistic. |
MartingaleState
dataclass
¶
MartingaleState(
step: int,
p_value: float,
log_martingale: float,
martingale: float,
log_restarted_martingale: float,
restarted_martingale: float,
log_cusum: float,
cusum: float,
log_shiryaev_roberts: float,
shiryaev_roberts: float,
triggered_alarms: tuple[str, ...],
log_e_value: float = 0.0,
e_value: float = 1.0,
)
Immutable snapshot of evidence statistics after one update.
Linear-scale values may be 0 or inf after floating-point underflow or
overflow; the corresponding log_* field preserves the log-scale state.
triggered_alarms reports thresholds crossed at this step and is not a
latched alarm history.
e_value is the ordinary capital ratio, not necessarily an increment
that reproduces the other statistics (see :class:MixtureMartingale).
BaseMartingale ¶
BaseMartingale(alarm_config: AlarmConfig | None = None)
Bases: ABC
Abstract base class for p-value-driven sequential evidence.
The default update multiplies capital by a non-negative betting factor and
updates alarm statistics from that factor. Composite subclasses may instead
aggregate component statistics. Ville-threshold interpretation
requires the resulting process to be an e-process under the null, which in
turn depends on the conditional validity of the input p-values. Merely
passing values in [0, 1] does not establish that property.
Source code in nonconform/martingales.py
198 199 200 | |
reset ¶
reset() -> None
Reset martingale and alarm statistics to initial values.
Source code in nonconform/martingales.py
207 208 209 210 211 212 213 214 215 216 217 218 | |
update_many ¶
update_many(
p_values: Sequence[float] | ndarray,
) -> list[MartingaleState]
Update state for each p-value in order and return all snapshots.
Source code in nonconform/martingales.py
220 221 222 223 224 | |
update ¶
update(p_value: float) -> MartingaleState
Ingest one p-value in [0, 1] and return the updated state.
Source code in nonconform/martingales.py
226 227 228 229 230 231 232 233 234 235 236 237 | |
PowerMartingale ¶
PowerMartingale(
epsilon: float = 0.5,
alarm_config: AlarmConfig | None = None,
)
Bases: BaseMartingale
Power martingale with fixed epsilon in (0, 1].
Each p-value contributes the betting factor
epsilon * p_value ** (epsilon - 1). Values below one emphasize small
p-values; epsilon=1 produces the constant factor one.
Source code in nonconform/martingales.py
321 322 323 324 325 326 327 328 329 | |
SimpleMixtureMartingale ¶
SimpleMixtureMartingale(
epsilons: Sequence[float] | ndarray | None = None,
*,
n_grid: int = 100,
min_epsilon: float = 0.01,
alarm_config: AlarmConfig | None = None,
)
Bases: BaseMartingale
Equal-weight mixture of power martingales over an epsilon grid.
When epsilons is omitted, the grid contains n_grid evenly spaced
values from min_epsilon through one. The mixture tracks the arithmetic
mean of component capitals, evaluated stably in log space.
Its alarm statistics use ratios of that all-history mixture capital. For
fixed-weight averages of independently maintained alarm statistics, use
:class:MixtureMartingale with separate :class:PowerMartingale experts.
Source code in nonconform/martingales.py
351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 | |
MixtureMartingale ¶
MixtureMartingale(
martingales: Sequence[BaseMartingale],
*,
weights: Sequence[float] | None = None,
alarm_config: AlarmConfig | None = None,
)
Bases: BaseMartingale
Fixed-weight averages of independently updated martingale evidence.
Each component receives the same p-value. Ordinary capital, harmonic restart evidence, CUSUM, and Shiryaev-Roberts statistics are averaged separately; alarm statistics are not computed from mixture capital ratios. Fixed normalized mixtures inherit the corresponding validity guarantees only when every participating component satisfies those guarantees under the common null and filtration. Composition does not remove learning inertia inside an individual expert.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
martingales
|
Sequence[BaseMartingale]
|
Nonempty sequence of reset |
required |
weights
|
Sequence[float] | None
|
Finite nonnegative weights with positive total. Normalized once; omitted weights give equal allocation. Zero-weight components are excluded from copying and updates. |
None
|
alarm_config
|
AlarmConfig | None
|
Thresholds applied to the combined statistics. Component alarms are not propagated. |
None
|
Notes
state.e_value and state.log_e_value describe the ordinary
mixture capital ratio only. They cannot reproduce the combined CUSUM,
SR, or restart statistics through a shared-increment recursion. For
consecutive zero capitals or consecutive infinite log-capitals, the
undefined ratio is reported as the neutral diagnostic factor one.
Finite log-capitals retain their ratio even when linear values overflow.
A component failure blocks further updates until reset(); the
mixture state remains at its last completed update. When used inside
an ExchangeabilityMonitor, reset the whole monitor after failure.
Aggregation adds O(J) work and storage beyond the J component costs.
Source code in nonconform/martingales.py
436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 | |
SimpleJumperMartingale ¶
SimpleJumperMartingale(
jump: float = 0.01,
alarm_config: AlarmConfig | None = None,
)
Bases: BaseMartingale
Simple Jumper martingale (Algorithm 1 in Vovk et al.).
This method mixes three betting components with parameters -1, 0,
and 1. Before each bet, jump redistributes capital equally among the
components; the current p-value then multiplies each component by
1 + epsilon * (p_value - 0.5).
Source code in nonconform/martingales.py
541 542 543 544 545 546 547 548 549 550 | |
Sequential Monitoring¶
nonconform.monitoring ¶
Sequential conformal monitoring with exact randomized ranks.
This module supplies the validity-critical first half of an exchangeability
martingale workflow: a frozen anomaly scoring rule followed by randomized
sequential ranks. The resulting p-values can be consumed by the betting
martingales in :mod:nonconform.martingales.
The existing :class:nonconform.ConformalDetector remains the batch and
pointwise conformal API. Its fixed-calibration p-values are deliberately not
modified by this module.
SequentialRankConformalizer ¶
SequentialRankConformalizer(
*, tail: Tail = "upper", seed: int | None = None
)
Generate exact randomized sequential conformal p-values.
At each update, the new score is ranked among all scores observed by this
object, including the new score itself. Independent uniform tie randomization
makes the sequential ranks independent Uniform(0, 1) variables when the
complete score sequence is exchangeable. Numeric input validation alone does
not establish exchangeability.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tail
|
Tail
|
|
'upper'
|
seed
|
int | None
|
Optional seed for a persistent random number generator. |
None
|
Notes
History grows without a sliding window. The current implementation uses an exact sorted list: rank queries are logarithmic, while insertion is linear in the history length.
Source code in nonconform/monitoring.py
124 125 126 127 128 129 | |
reset ¶
reset() -> None
Clear score history and restore the initial RNG state.
Source code in nonconform/monitoring.py
141 142 143 144 | |
prime ¶
prime(score: float) -> Self
Add one score to rank history without producing a p-value.
Priming does not update evidence and does not consume a random draw.
Source code in nonconform/monitoring.py
146 147 148 149 150 151 152 | |
prime_many ¶
prime_many(scores: Any) -> Self
Add a one-dimensional score collection without consuming RNG draws.
Source code in nonconform/monitoring.py
154 155 156 157 158 159 160 | |
update ¶
update(score: float) -> float
Insert one score and return its randomized sequential p-value.
The result lies in [0, 1) because the randomized rank uses a uniform
draw on a half-open interval.
Source code in nonconform/monitoring.py
162 163 164 165 166 167 168 | |
update_many ¶
update_many(scores: Any) -> np.ndarray
Process a one-dimensional score sequence in order.
Source code in nonconform/monitoring.py
188 189 190 191 192 193 | |
MonitorState
dataclass
¶
MonitorState(
rank_step: int,
score: float,
martingale_state: MartingaleState,
)
Immutable snapshot from one sequential monitoring update.
evidence_step
property
¶
evidence_step: int
Number of observations included in the evidence process.
e_value
property
¶
e_value: float
Ordinary capital ratio; mixtures aggregate alarm statistics separately.
restarted_martingale
property
¶
restarted_martingale: float
Harmonic restart-mixture e-process value.
triggered_alarms
property
¶
triggered_alarms: tuple[str, ...]
Alarm names whose configured thresholds are currently crossed.
ExchangeabilityMonitor ¶
ExchangeabilityMonitor(
detector: Any,
*,
conformalizer: SequentialRankConformalizer
| None = None,
martingale: BaseMartingale | None = None,
score_polarity: ScorePolarityInput = None,
seed: int | None = None,
)
Monitor a stream using a frozen scorer and sequential conformal ranks.
The scorer is fitted once and then applied row-wise to reference and stream
observations. prime establishes rank history without starting
evidence; update generates a randomized sequential p-value and sends it
to the configured martingale.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detector
|
Any
|
PyOD, scikit-learn, or custom anomaly detector. |
required |
conformalizer
|
SequentialRankConformalizer | None
|
Stateful sequential rank conformalizer. Defaults to an
upper-tail :class: |
None
|
martingale
|
BaseMartingale | None
|
P-value betting martingale. Defaults to Simple Jumper. |
None
|
score_polarity
|
ScorePolarityInput
|
Detector score direction, following
:class: |
None
|
seed
|
int | None
|
Seed propagated to the scorer and default conformalizer. |
None
|
Supplied conformalizers and martingales are treated as configuration prototypes and deep-copied. The monitor exclusively owns and mutates its scorer, rank history, and evidence process.
Notes
Sequential-rank validity requires the priming and monitored scores to be exchangeable under the null, conditional on a fixed training-only scoring construction. Ville guarantees additionally require a valid e-process. Do not refit the scorer during an episode.
Source code in nonconform/monitoring.py
268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 | |
state
property
¶
state: MonitorState | None
Return the latest monitoring state, or None before the first update.
from_split_detector
classmethod
¶
from_split_detector(
detector: ConformalDetector,
*,
conformalizer: SequentialRankConformalizer
| None = None,
martingale: BaseMartingale | None = None,
seed: int | None = None,
) -> ExchangeabilityMonitor
Create a monitor from a fitted unweighted Split detector.
The fitted scoring model is copied and frozen. Existing calibration scores initialize sequential rank history; they are not reused as a fixed empirical CDF and do not contribute to martingale capital.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detector
|
ConformalDetector
|
Fitted, unweighted detector using |
required |
conformalizer
|
SequentialRankConformalizer | None
|
Optional empty conformalizer configuration prototype. |
None
|
martingale
|
BaseMartingale | None
|
Optional reset martingale configuration prototype. |
None
|
seed
|
int | None
|
Seed for the monitor's conformalizer, or None. |
None
|
Returns:
| Type | Description |
|---|---|
ExchangeabilityMonitor
|
A fitted monitor primed with copied calibration scores. |
Source code in nonconform/monitoring.py
334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 | |
fit ¶
fit(x: DataFrame | ndarray) -> Self
Fit the scoring rule once on a proper training set.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
DataFrame | ndarray
|
Finite two-dimensional fitting data. It must be independent of the later null rank sequence for the documented validity scope. |
required |
Returns:
| Type | Description |
|---|---|
Self
|
This monitor with empty rank and evidence state. |
Source code in nonconform/monitoring.py
382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 | |
reset ¶
reset() -> None
Reset rank and evidence state while retaining the fitted scorer.
A reset starts a new monitoring episode. Repeated episodes need their own error-budget accounting if a lifetime false-alarm guarantee is required.
Source code in nonconform/monitoring.py
399 400 401 402 403 404 405 406 407 408 | |
prime ¶
prime(x: DataFrame | ndarray) -> Self
Score reference observations and add them only to rank history.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
DataFrame | ndarray
|
Finite two-dimensional reference batch with the fitted feature count. |
required |
Returns:
| Type | Description |
|---|---|
Self
|
This monitor after extending its rank history. |
Source code in nonconform/monitoring.py
410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 | |
update ¶
update(x: Series | ndarray) -> MonitorState
Score one observation and update sequential evidence.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
Series | ndarray
|
One finite feature vector with the fitted feature count. |
required |
Returns:
| Type | Description |
|---|---|
MonitorState
|
The immutable state after ranking, betting, and alarm evaluation. |
Source code in nonconform/monitoring.py
430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 | |
update_many ¶
update_many(x: DataFrame | ndarray) -> list[MonitorState]
Update sequentially for every row in input order.
The input is accepted as a two-dimensional batch for convenience, but the method preserves row order and performs one sequential update per row.
Source code in nonconform/monitoring.py
458 459 460 461 462 463 464 465 466 467 | |
Metrics¶
nonconform.metrics ¶
Public helpers for score aggregation and labeled evaluation.
false_discovery_rate retains its v1 public name but returns the realized
false discovery proportion for one supplied testing family. statistical_power
similarly returns the realized true positive rate. Expected FDR and statistical
power are repeated-sampling properties, not quantities identified by one labeled
family.
aggregate ¶
aggregate(method: str, scores: ndarray) -> np.ndarray
Aggregate anomaly scores using a specified method.
Applies a chosen aggregation technique to a 2D array of anomaly scores, where each row represents scores from a different model and each column corresponds to a data sample.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
method
|
str
|
The aggregation method to apply. |
required |
scores
|
ndarray
|
A 2D array of anomaly scores. Rows = different models, columns = data samples. Aggregation is performed along axis=0. |
required |
Returns:
| Type | Description |
|---|---|
ndarray
|
Array of aggregated anomaly scores with length equal to number |
ndarray
|
of columns in input. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the method is not a supported aggregation type. |
Examples:
>>> import numpy as np
>>> from nonconform.metrics import aggregate
>>> scores = np.array([[1, 2, 3], [4, 5, 6]])
>>> aggregate("mean", scores)
array([2.5, 3.5, 4.5])
Source code in nonconform/_internal/math_utils.py
45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 | |
false_discovery_rate ¶
false_discovery_rate(y: ndarray, y_hat: ndarray) -> float
Calculate the realized false discovery proportion for one labeled family.
The returned quantity is FP / (FP + TP) for the supplied realization.
Its expectation over repeated testing families is the false discovery rate
(FDR). The function name is retained for public API compatibility.
If there are no predicted positives, this function returns a realized FDP of 0.0.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
y
|
ndarray
|
True binary labels (1 = positive/anomaly, 0 = negative/normal). |
required |
y_hat
|
ndarray
|
Predicted binary labels. |
required |
Returns:
| Type | Description |
|---|---|
float
|
The realized false discovery proportion. |
Examples:
>>> import numpy as np
>>> from nonconform.metrics import false_discovery_rate
>>> y = np.array([1, 0, 1, 0])
>>> y_hat = np.array([1, 1, 0, 0]) # 1 TP, 1 FP
>>> float(false_discovery_rate(y, y_hat))
0.5
Source code in nonconform/_internal/math_utils.py
84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 | |
statistical_power ¶
statistical_power(y: ndarray, y_hat: ndarray) -> float
Calculate realized recall (true positive rate) for one labeled family.
The returned quantity is TP / (TP + FN) for the supplied realization.
Its expectation under a specified data-generating process is statistical
power. The function name is retained for public API compatibility.
If there are no actual positives, this function returns a realized true positive rate of 0.0.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
y
|
ndarray
|
True binary labels (1 = positive/anomaly, 0 = negative/normal). |
required |
y_hat
|
ndarray
|
Predicted binary labels. |
required |
Returns:
| Type | Description |
|---|---|
float
|
The realized true positive rate. |
Examples:
>>> import numpy as np
>>> from nonconform.metrics import statistical_power
>>> y = np.array([1, 0, 1, 0])
>>> y_hat = np.array([1, 1, 0, 0]) # 1 TP, 1 FN
>>> float(statistical_power(y, y_hat))
0.5
Source code in nonconform/_internal/math_utils.py
123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 | |
Enumerations¶
nonconform.enums ¶
Public enumeration constants for nonconform.
ConformalMode ¶
Bases: Enum
Model retention modes for conformal resampling strategies.
Attributes:
| Name | Type | Description |
|---|---|---|
PLUS |
Retain calibration-time models and aggregate their test scores. |
|
SINGLE_MODEL |
Fit and retain one final model after constructing calibration scores. |
Distribution ¶
Bases: Enum
Reserved distribution choices for randomized size configurations.
The enum remains part of the v1 public surface through nonconform.enums.
Current public calibration strategy constructors do not consume it.
Attributes:
| Name | Type | Description |
|---|---|---|
BETA_BINOMIAL |
Beta-binomial distribution for drawing validation fractions. |
|
UNIFORM |
Discrete uniform distribution over a specified range. |
|
GRID |
Discrete distribution over a specified set of values. |
Kernel ¶
Bases: Enum
Kernel functions for KDE-based score-tail estimation.
Attributes:
| Name | Type | Description |
|---|---|---|
GAUSSIAN |
Gaussian (normal) kernel. |
|
EXPONENTIAL |
Exponential kernel. |
|
BOX |
Box (uniform) kernel. |
|
TRIANGULAR |
Triangular kernel. |
|
EPANECHNIKOV |
Epanechnikov kernel. |
|
BIWEIGHT |
Biweight (quartic) kernel. |
|
TRIWEIGHT |
Triweight kernel. |
|
TRICUBE |
Tricube kernel. |
|
COSINE |
Cosine kernel. |
Pruning ¶
Bases: Enum
Pruning strategies for weighted FDR control.
Attributes:
| Name | Type | Description |
|---|---|---|
HETEROGENEOUS |
Use independent uniform draws for candidate-specific randomized WCS pruning. |
|
HOMOGENEOUS |
Use one shared uniform draw for randomized WCS pruning. |
|
DETERMINISTIC |
Use the non-randomized WCS pruning rule. |
ScorePolarity ¶
Bases: Enum
Score direction conventions for anomaly detectors.
Attributes:
| Name | Type | Description |
|---|---|---|
AUTO |
Strictly infer polarity from recognized detector families and raise for an unknown custom detector. |
|
HIGHER_IS_ANOMALOUS |
Higher scores indicate more anomalous samples. |
|
HIGHER_IS_NORMAL |
Higher scores indicate more normal samples. |
TieBreakMode ¶
Bases: Enum
Tie-breaking modes for empirical p-value estimation.
Attributes:
| Name | Type | Description |
|---|---|---|
CLASSICAL |
Deterministic empirical conformal formula. |
|
RANDOMIZED |
Randomized smoothing with uniform tie-breaking. |
Data Structures¶
nonconform.structures ¶
Core data structures and protocols for nonconform.
This module provides the fundamental types used throughout the package:
Classes:
| Name | Description |
|---|---|
AnomalyDetector |
Protocol defining the detector interface. |
ConformalResult |
Container for conformal inference outputs. |
AnomalyDetector ¶
Bases: Protocol
Protocol defining the interface for anomaly detectors.
A supported PyOD model, a recognized scikit-learn estimator, or a custom object can be used with nonconform when it provides this interface. The object must also support shallow and deep copying because resampling strategies create independent detector replicas.
Required methods
fit: Train the detector on data decision_function: Compute anomaly scores get_params: Retrieve detector parameters set_params: Configure detector parameters
Examples:
from sklearn.ensemble import IsolationForest
from nonconform.structures import AnomalyDetector
detector: AnomalyDetector = IsolationForest(random_state=42)
print(isinstance(detector, AnomalyDetector))
fit ¶
fit(X: ndarray, y: ndarray | None = None) -> Self
Train the anomaly detector.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
ndarray
|
Training data of shape (n_samples, n_features). |
required |
y
|
ndarray | None
|
Ignored. Present for API consistency. |
None
|
Returns:
| Type | Description |
|---|---|
Self
|
The fitted detector instance. |
Source code in nonconform/structures.py
56 57 58 59 60 61 62 63 64 65 66 | |
decision_function ¶
decision_function(X: ndarray) -> np.ndarray
Compute anomaly scores for samples.
Score direction is detector-specific. Pass the corresponding
score_polarity to :class:~nonconform.detector.ConformalDetector
when it cannot be inferred safely.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
ndarray
|
Data of shape (n_samples, n_features). |
required |
Returns:
| Type | Description |
|---|---|
ndarray
|
Anomaly scores of shape (n_samples,). |
Source code in nonconform/structures.py
68 69 70 71 72 73 74 75 76 77 78 79 80 81 | |
get_params ¶
get_params(deep: bool = True) -> dict[str, Any]
Get parameters for this detector.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
deep
|
bool
|
If True, return parameters for sub-objects. |
True
|
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
Parameter names mapped to their values. |
Source code in nonconform/structures.py
83 84 85 86 87 88 89 90 91 92 | |
set_params ¶
set_params(**params: Any) -> Self
Set parameters for this detector.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
**params
|
Any
|
Detector parameters. |
{}
|
Returns:
| Type | Description |
|---|---|
Self
|
The detector instance. |
Source code in nonconform/structures.py
94 95 96 97 98 99 100 101 102 103 | |
ConformalResult
dataclass
¶
ConformalResult(
p_values: ndarray | None = None,
test_scores: ndarray | None = None,
calib_scores: ndarray | None = None,
test_weights: ndarray | None = None,
calib_weights: ndarray | None = None,
metadata: dict[str, Any] = dict(),
)
Snapshot of detector outputs for downstream procedures.
This dataclass holds the latest p-values or score-tail estimates, raw scores, and optional weights produced by a detector call.
Attributes:
| Name | Type | Description |
|---|---|---|
p_values |
ndarray | None
|
Values produced by the configured estimation strategy, or None
when only scores were requested. With |
test_scores |
ndarray | None
|
Aggregated, anomalous-higher scores for test instances. |
calib_scores |
ndarray | None
|
Anomalous-higher scores for the calibration set, or None for aggregated raw DerandomizedSplits scores, which have no single corresponding calibration distribution. |
test_weights |
ndarray | None
|
Importance weights for test instances (weighted mode only). |
calib_weights |
ndarray | None
|
Importance weights for calibration instances. |
metadata |
dict[str, Any]
|
Method metadata, including the strategy, estimator, and weighted-mode marker for p-value computations. |
Examples:
import numpy as np
from nonconform.structures import ConformalResult
result = ConformalResult(
p_values=np.array([0.50, 0.02]),
test_scores=np.array([0.1, 2.4]),
calib_scores=np.array([-0.2, 0.0, 0.3, 0.8]),
metadata={"nonconform": {"weighted": False}},
)
print(result.p_values)
print(result.metadata)
fdp_bounds ¶
fdp_bounds(
*,
confidence: float = 0.95,
method: str = "mc_thc",
n_resamples: int | None = None,
seed: int | None = None,
boost: bool = True,
lower: float | None = None,
upper: float | None = None,
beta: float | None = None,
precision: float | None = None,
) -> FDPCertificate
Certify this snapshot without rescoring.
Requires an unmodified native snapshot of unweighted empirical Split inference, including detached calibration. Scope and batch dimensions are checked; these checks do not prove array integrity or scientific exchangeability. For external p-values use FDPCertificate.from_p_values.
Options and defaults match :meth:nonconform.fdr.FDPCertificate.from_p_values.
Confidence is simultaneous coverage, not an FDR target. The Monte Carlo
seed is independent of the detector seed. Query cutoffs on the returned
immutable certificate; later edits to this snapshot cannot affect it.
Source code in nonconform/structures.py
173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 | |
copy ¶
copy() -> ConformalResult
Return a copy with arrays and metadata fully duplicated.
Returns:
| Type | Description |
|---|---|
ConformalResult
|
A new ConformalResult with copied arrays and deep-copied metadata. |
Source code in nonconform/structures.py
216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 | |
Adapters¶
nonconform.adapters ¶
External detector adapters for nonconform.
ScorePolarityAdapter ¶
ScorePolarityAdapter(
detector: AnomalyDetector, score_polarity: ScorePolarity
)
Normalize a wrapped detector to anomalous-higher scores.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detector
|
AnomalyDetector
|
Detector implementing :class: |
required |
score_polarity
|
ScorePolarity
|
Convention of its raw scores. Automatic polarity is not accepted here. |
required |
Source code in nonconform/adapters.py
277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 | |
fit ¶
fit(X: ndarray, y: ndarray | None = None) -> Self
Fit wrapped detector.
Source code in nonconform/adapters.py
295 296 297 298 | |
decision_function ¶
decision_function(X: ndarray) -> np.ndarray
Return scores transformed to anomalous-higher convention.
Source code in nonconform/adapters.py
300 301 302 303 | |
get_params ¶
get_params(deep: bool = True) -> dict[str, Any]
Delegate parameter retrieval to wrapped detector.
Source code in nonconform/adapters.py
305 306 307 | |
set_params ¶
set_params(**params: Any) -> Self
Delegate parameter updates to wrapped detector.
Source code in nonconform/adapters.py
309 310 311 312 | |
PyODAdapter ¶
PyODAdapter(detector: Any)
Wrap a PyOD detector with the public anomaly-detector protocol.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detector
|
Any
|
Configured PyOD detector. |
required |
Raises:
| Type | Description |
|---|---|
ImportError
|
If PyOD is not installed. |
Source code in nonconform/adapters.py
353 354 355 356 357 | |
fit ¶
fit(X: ndarray, y: ndarray | None = None) -> Self
Fit wrapped detector.
Source code in nonconform/adapters.py
359 360 361 362 | |
decision_function ¶
decision_function(X: ndarray) -> np.ndarray
Return anomaly scores from wrapped detector.
Source code in nonconform/adapters.py
364 365 366 | |
get_params ¶
get_params(deep: bool = True) -> dict[str, Any]
Delegate parameter retrieval to wrapped detector.
Source code in nonconform/adapters.py
368 369 370 | |
set_params ¶
set_params(**params: Any) -> Self
Delegate parameter updates to wrapped detector.
Source code in nonconform/adapters.py
372 373 374 375 | |
adapt ¶
adapt(detector: Any) -> AnomalyDetector
Return a detector that satisfies the public anomaly-detector protocol.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detector
|
Any
|
PyOD, scikit-learn, or custom detector object. |
required |
Returns:
| Type | Description |
|---|---|
AnomalyDetector
|
The original structurally compatible object or a PyOD adapter. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the detector is a blocked batch-adaptive PyOD class. |
ImportError
|
If the object appears to require PyOD but PyOD is absent. |
TypeError
|
If required protocol methods are missing. |
Source code in nonconform/adapters.py
71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 | |
parse_score_polarity ¶
parse_score_polarity(
score_polarity: ScorePolarityInput,
) -> ScorePolarity
Normalize a score-polarity string or enum.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
score_polarity
|
ScorePolarityInput
|
A :class: |
required |
Returns:
| Type | Description |
|---|---|
ScorePolarity
|
The corresponding :class: |
Raises:
| Type | Description |
|---|---|
ValueError
|
If a string value is unsupported. |
TypeError
|
If the input is neither a string nor |
Source code in nonconform/adapters.py
110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 | |
resolve_implicit_score_polarity ¶
resolve_implicit_score_polarity(
detector: Any,
) -> ScorePolarity
Resolve score polarity when users omit score_polarity.
The default favors low-friction custom detector onboarding while preserving safe behavior for known detector families: - Known sklearn normality detectors -> HIGHER_IS_NORMAL - PyOD detectors -> HIGHER_IS_ANOMALOUS - Unknown custom detectors -> HIGHER_IS_ANOMALOUS
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detector
|
Any
|
Adapted detector whose family determines the default. |
required |
Returns:
| Type | Description |
|---|---|
ScorePolarity
|
The anomalous-higher or normal-higher convention selected by the v1 |
ScorePolarity
|
implicit-default policy. |
Source code in nonconform/adapters.py
169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 | |
resolve_score_polarity ¶
resolve_score_polarity(
detector: Any, score_polarity: ScorePolarityInput
) -> ScorePolarity
Resolve requested score polarity in strict AUTO mode.
Unlike resolve_implicit_score_polarity, this function is intentionally
strict for explicit score_polarity="auto" and raises for unknown
detectors.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detector
|
Any
|
Adapted detector whose family may be inspected. |
required |
score_polarity
|
ScorePolarityInput
|
Explicit score convention or |
required |
Returns:
| Type | Description |
|---|---|
ScorePolarity
|
A non-auto score-polarity convention. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If strict automatic inference cannot identify the detector family, or if the value is unsupported. |
TypeError
|
If |
Source code in nonconform/adapters.py
192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 | |
apply_score_polarity ¶
apply_score_polarity(
detector: AnomalyDetector,
score_polarity: ScorePolarityInput,
) -> AnomalyDetector
Normalize detector output so larger values mean more anomalous.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
detector
|
AnomalyDetector
|
Detector whose raw score convention is known. |
required |
score_polarity
|
ScorePolarityInput
|
Convention of the detector's raw scores. |
required |
Returns:
| Type | Description |
|---|---|
AnomalyDetector
|
The original detector for anomalous-higher scores, or a polarity adapter |
AnomalyDetector
|
that negates normal-higher scores. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If unresolved automatic polarity is supplied. |
Source code in nonconform/adapters.py
234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 | |