backup: 2026-07-25 14:39

This commit is contained in:
2026-07-25 14:39:02 +03:00
parent 3a2b6bfde8
commit 3e65bf2dab
19 changed files with 884 additions and 32 deletions
+54
View File
@@ -4,3 +4,57 @@
#+title: impl/loss-functions
#+filetags: :project: :knowledge: :loss-functions:
* impl/loss-functions
** KernelizedROLLoss — Key Design Decisions
- =_icdf= uses **bisection** (not Newton-Raphson). NR was replaced because =F_prime= goes near-zero when =v= is large and =tau= drifts, causing divergence. Bisection bracket is =[min(scores)10/v, max(scores)+10/v]=, guaranteed to converge.
- Score normalisation (=normalize= flag): divides =yh= by its std before KDE. Can be disabled via =kernelized_roll_fpr(fpr, normalize=False)=. The chain-rule correction (multiply grad by =score_scale=) is applied in backward.
- =_grad_tau_scores_neg= denominator clamped at =1e-10= to prevent 0/0 when all sigmoid kernels saturate.
- Variance clamped at =1e-8= in normalisation to prevent =inf= when all batch scores are identical.
** KernelizedROLLoss Unification (2026-07-14)
=KernelizedROLLoss= and =KernelizedROLLossAOC= merged into a single class accepting =alphas= (list of CDF targets) and =divisor=:
- Point-FPR: =KernelizedROLLoss.apply([1fpr], 1, yh, y, gamma, normalize)=
- AOC: =KernelizedROLLoss.apply(_AOC_ALPHAS, _AOC_DIVISOR, yh, y, gamma, normalize)=
Backward always vectorised (K-dim); K=1 reduces to the scalar case. =ctx.divisor= stores the non-tensor divisor.
** BCE Combination
Both =kernelized_roll_fpr= and =kernelized_roll_tpr= accept =bce_weight= (default 0.0) and =bce_pos_weight= (default 1.0).
Loss = ROLL + bce_weight × BCE. Standard value: =bce_weight=0.5=, =bce_pos_weight=num_false/num_true=.
** ISJ Bandwidth Failure
ISJ fails when a class has too few samples in a split (~<20). Affects: cleveland, ecoli, glass4, glass5.
** Gamma Scheduling
=KernelScheduler= widens KDE kernels early then anneals. Default: =initial_gamma=100=, =decay=0.5=, =decay_every=500=. Gamma is a *divisor* of =v=.
** LibAUC Integration (2026-07-14, tuned 2026-07-14)
=libauc_auc_loss= in =src/roll.py= wraps =AUCMLoss= (Yuan et al., ICCV 2021) to the =(yh, y)= convention.
- Applies =torch.sigmoid(yh)= before passing to AUCMLoss (squared-hinge surrogate, no built-in sigmoid).
- AUCMLoss has learnable params =a, b, alpha= — must use PESG optimizer (not Adam).
- =opt_factory= signature: =opt_factory(model)= — does NOT take criteriorator (would be unbound at call site in =_perform_episode=).
- =opt_factory= pattern: =loss_fn.pesg_opt_factory(...)= returns picklable =_PESGFactory= sharing the same inner =AUCMLoss= instance as the criteriorator's loss func (preserved by pickle memo).
- =ExperimentConfiguration= gained =opt_factory: callable = None=; overrides =optim_class/optim_args= when set.
- =tqdm= is an undeclared dep of libauc — must be in =propagatedBuildInputs= in flake.nix.
*** Recommended settings (v2 + adam, tuned 2026-07-14)
- =version='v2'= — removes class prior from formulation; LibAUC explicitly recommends over v1; default in =libauc_auc_loss=.
- =mode='adam'= in PESG — Adam-style primal updates, converges much faster than SGD on small datasets; default in =pesg_opt_factory=.
- =imratio= passed explicitly as =num_true / (num_true + num_false)= — avoids noisy mini-batch estimates on imbalanced data. Note: v2 ignores imratio (prior-free), so this only matters if reverting to v1.
- =lr=0.1= (PESG default, much larger than Adam's 1e-3). If diverging: try =epoch_decay=0.0= or drop lr to =0.01=.
** Grace Period (BasicCriteriorator / CRBasedCriteriorator, 2026-07-14)
=grace_period: int = 0= parameter added to both criteriorators. During the first =grace_period= epochs:
- =best_flag= is always False (no model snapshot taken)
- Patience counter does not increment (early stopping cannot fire)
All keel configs use =grace_period=50=. Prevents noisy early-epoch models from being selected as best.