Analyzing Non‐Random Selectivity in Online Job Advertisements Using Eurostat Benchmark Data and Generalized Sample Selection Models: An Application to <scp>EU</scp> Regional Labor Markets
Pietro Giorgio Lovaglio & Mario Mezzanzanica
What the paper says
ABSTRACT The present paper provides an overall framework to afford the problem of non‐representativeness and non‐random selectivity arising from online job ads data, using Generalized sample selection models and Eurostat benchmark data. We jointly model the outcome intensity (number of online job ads in observed profiles, whose levels are defined by auxiliary variables) and the probability of endogenous selection (likelihood that online job ads are not missing in a given profile), allowing us to model the missing data mechanism without the need of a priori justification of missingness at random, as generally supposed by multilevel regression and post‐stratification, a popular benchmark technique in this field. Moreover, we offer new post‐stratification strategies to calibrate the unconditional predictions on benchmark/reference samples. We use data from the Cedefop's Skill Ovate platform collecting online job advertisements for all EU regions in 2022 and an Italian web‐platform during 2013Q2‐2018Q2, whereas as reference samples, aggregated LFS recent job starters and LFS new hires from microdata that represent reasonable lower bounds for job advertisements. Online job ads present a strong overrepresentation with respect to benchmark data (+40% with respect to LFS recent job starters and +400% over new hires from LFS microdata), whereas generalized sample selection models reduced this bias by half, unlike Multilevel post‐stratification and other univariate approaches, which furthermore resulted in bias.
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.50 × 0.4 = 0.20 |
| M · momentum | 0.50 × 0.15 = 0.07 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.