Testing Hypotheses of Covariate Effects on Topics of Discourse
Gabriel Phelan & David A. Campbell
What the paper says
We introduce an approach to topic modeling with document‐level covariates that remains tractable in the face of large text corpora. This is achieved by de‐emphasizing the role of parameter estimation in an underlying probabilistic model, assuming instead that the data come from a fixed but unknown distribution whose statistical functionals are of interest. We propose combining a convex formulation of non‐negative matrix factorization with standard regression techniques as a fast‐to‐compute and useful estimate of such a functional. Uncertainty quantification can then be achieved by reposing non‐parametric resampling methods on top of this scheme. This is in contrast to popular topic modeling paradigms, which posit a complex and often hard‐to‐fit generative model of the data. We argue that the simple, non‐parametric approach advocated here is faster, more interpretable, and enjoys better inferential justification than said generative models. Finally, our methods are demonstrated with an application analyzing covariate effects on discourse of flavors attributed to Canadian beers.
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.50 × 0.4 = 0.20 |
| M · momentum | 0.50 × 0.15 = 0.07 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.