From raw text to fairseq RoBERTa: A modular snakemake-based framework enabling language-specific BPE tokenization

Raphael Schmitt

Software Impacts2026https://doi.org/10.1016/j.simpa.2026.100824article
AJG 1
Weight
0.50

What the paper says

Large-scale language model training requires robust and reproducible data preprocessing. While fairseq provides efficient training routines for RoBERTa models, preparing high-quality, language-specific data remains complex. We present modular Snakemake-based workflows for large-scale language model preparation, covering filtering, GPT-2 BPE tokenization, and fairseq-compatible data generation. The pipelines support new and existing tokenizers, enable scalable HPC parallelism, and include utilities for converting trained models to the Huggingface format. Bundled with a fairseq fork supporting GPU clusters and Cloud TPUs, the framework has been used to train GottBERT, GeistBERT, ChristBERT, PortBERT, SindBERT, and HalleluBERT, and generalizes into a reusable preprocessing infrastructure. • Modular Snakemake framework for large-scale RoBERTa pre-processing. • Supports high-precision corpus filtering and language-specific BPE. • Includes a fairseq fork with TPU v3/v4 support and Whole Word Masking (Huggingface tokenizers on GPU). • Provides utilities for Huggingface conversion and log monitoring. • Applied to pre-process and train GottBERT, GeistBERT, ChristBERT, PortBERT, SindBERT and HalleluBERT.

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.1016/j.simpa.2026.100824

Or copy a formatted citation

@article{raphael2026,
  title        = {{From raw text to fairseq RoBERTa: A modular snakemake-based framework enabling language-specific BPE tokenization}},
  author       = {Raphael Schmitt},
  journal      = {Software Impacts},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.1016/j.simpa.2026.100824},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

From raw text to fairseq RoBERTa: A modular snakemake-based framework enabling language-specific BPE tokenization

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.