Pseudo-Labeling for kernel ridge regression under covariate shift

Kaizheng Wang

Annals of Statistics2026https://doi.org/10.1214/25-aos2566article
AJG 4*ABDC A*
Weight
0.50

What the paper says

We develop and analyze a principled approach to kernel ridge regression under covariate shift. The goal is to learn a regression function with small mean squared error over a target distribution, based on unlabeled data from there and labeled data that may have a different feature distribution. We propose to split the labeled data into two subsets, and conduct kernel ridge regression on them separately to obtain a collection of candidate models and an imputation model. We use the latter to fill the missing labels and then select the best candidate accordingly. Our nonasymptotic excess risk bounds demonstrate that our estimator adapts effectively to both the structure of the target distribution and the covariate shift. This adaptation is quantified through a notion of effective sample size that reflects the value of labeled source data for the target regression task. Our estimator achieves the minimax optimal error rate up to a polylogarithmic factor, and we find that using pseudo-labels for model selection does not significantly hinder performance.

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.1214/25-aos2566

Or copy a formatted citation

@article{kaizheng2026,
  title        = {{Pseudo-Labeling for kernel ridge regression under covariate shift}},
  author       = {Kaizheng Wang},
  journal      = {Annals of Statistics},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.1214/25-aos2566},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

Pseudo-Labeling for kernel ridge regression under covariate shift

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.