Split-Apply-Combine with Dynamic Grouping

Mark P. J. van der Loo

Journal of Statistical Software2025https://doi.org/10.18637/jss.v112.i04article
ABDC A
Weight
0.37

What the paper says

Partitioning a data set by one or more of its attributes and computing an aggregate for each part is one of the most common operations in data analyses. There are use cases where the partitioning is determined dynamically by collapsing smaller subsets into larger ones, to ensure sufficient support for the computed aggregate. These use cases are not supported by software implementing split-apply-combine types of operations. This paper presents the R package accumulate that offers convenient interfaces for defining grouped aggregation where the grouping itself is dynamically determined, based on user-defined conditions on subsets, and a user-defined subset collapsing scheme. The formal underlying algorithm is described and analyzed as well.

1 citation

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.18637/jss.v112.i04

Or copy a formatted citation

@article{mark2025,
  title        = {{Split-Apply-Combine with Dynamic Grouping}},
  author       = {Mark P. J. van der Loo},
  journal      = {Journal of Statistical Software},
  year         = {2025},
  doi          = {https://doi.org/https://doi.org/10.18637/jss.v112.i04},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

Split-Apply-Combine with Dynamic Grouping

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.37

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.16 × 0.4 = 0.06
M · momentum0.53 × 0.15 = 0.08
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.