Research
TinyFabulist
TinyFabulist is a three-stage research program built around controlled synthetic narratives.
TF1 — generation
Structured six-slot specifications are expanded into prompts and passed to open-weight generators no larger than 8B parameters. The released artifact is a three-million-story English corpus with generation metadata and evaluation signals.
- Paper: TF1-EN-3M
- The paper records the dataset schema, generation pipeline, and public release.
TF2 — English–Romanian translation
TF2 turns selected TF1 stories into English–Romanian parallel resources and tests how far a fine-tuned open model can narrow the gap to much larger systems at lower inference cost. The project separates the 15K reference set, the three-million-pair corpus, and the released model; they serve different purposes.
- Journal article: Building Large-Scale English–Romanian Literary Translation Resources with Open Models
- Versioned preprint
- The article distinguishes the full three-million-pair corpus from the curated 15K tuning and evaluation set.
TF3 — compact Romanian models
TF3 covers tokenizer construction, from-scratch pre-training, compression, evaluation, and Romanian-native generation. The current paper distinguishes the 51.65M-parameter model from the 26.45M-parameter compressed student; neither number should be reduced to “a 50M model” without context.
Evaluation and Romanian NLP
Two supporting lines of work cut across the pipeline:
- a survey of synthetic text and code generation (arXiv);
- Romanian diacritic restoration, including the InnoComp 2025 study published in Springer CCIS 2794 (paper, preprint).
The open-weight panel now has one measured use: under the fixed TF1 protocol, its system ranking tracked o4-mini at Spearman (\rho=0.93) and Kendall (\tau=0.78). Item-level agreement remained weak, and the planned human arbitration has not run. I therefore use the panel for aggregate comparison under the tested protocol, not as a replacement for human review or as an automatic filter for individual texts.
From the thesis draft
The current proposal-stage thesis joins the three systems under a measurement question: when a pipeline is called controlled, which output properties were actually tested? This series develops the results that are not already covered by the release notes.
-
Three Million Stories Are Still a Sparse Sample — What a three-million-row corpus covers when six slots each have one hundred possible values.
-
Change One Slot, Watch What Else Moves — A paired intervention test separates response to a requested field from leakage into a field that should stay fixed.
-
The Interval Belongs to the Comparison — Why two model gaps require different bootstrap designs—and why one missing record prevents an interval altogether.
-
The Panel Was Weak on Items and Useful for Ranking Systems — An open-weight judge panel was weak at item-level agreement yet useful for ranking systems in one fixed protocol.
-
The 2.4M-Parameter Model Won—Until the Text Got Noisy — Romanian diacritic restoration changes winners when clean benchmark text gives way to typos and OCR-like corruption.
-
Ten Thousand Adaptation Steps Could Not Remove One Scraper Artifact — One Romanian checkpoint kept emitting a web-page token after supervised adaptation, turning contamination into a model-selection failure.
Reproducibility
Public papers, datasets, prompts, and code improve auditability. They do not make hardware behavior, third-party model serving, or stochastic execution automatically identical. Each technical note records the boundary it actually tested.