TF2: Open Models for English–Romanian Literary Translation
The model family, 15K silver-reference set, and three-million-pair corpus released with the TF2 preprint.
TF2 extends TinyFabulist from English generation into English-Romanian literary translation. The release includes tuning data, adapted open models, evaluation records, and a large translated corpus, but those artifacts were produced at different stages and should not be treated as one dataset.
The 15K and three-million-pair releases must not be called one dataset. The first is a controlled set of synthetic silver references; the second is the scale artifact. They answer different questions and carry different guarantees.
Version note. This retrospective was published here in August 2026 and follows the v4 preprint and Frontiers article, not only the September 2025 submission. The paper’s current title is Building Large-Scale English–Romanian Literary Translation Resources with Open Models.
Two sets, two jobs
The smaller set pairs 15,000 English fables with Romanian references generated by GPT-o3, split 12,000/1,500/1,500 for training, validation, and testing. It supports instruction tuning and a controlled evaluation target. “Silver” matters here: these are model-generated references, not human translations.
The three-million-pair corpus is built for scale. It is useful for downstream training and analysis, but it is not interchangeable with either the 15K set or a human-authored literary parallel corpus. The release also includes fine-tuned 1B, 4B, and 12B models; the 12B result is the main comparison reported below.
Training and evaluation
The 12B model uses LoRA with rank 32, scaling factor 32, and 0.05 dropout across the attention projections and MLP projections. That configuration is worth recording because “fine-tuned” otherwise hides most of the reproducible decision.
On the held-out set, the base 12B model scored 4.43 under the five-part judge rubric; TF2-12B scored 4.83. Corpus BLEU moved from 0.0214 to 0.0926. The quantized model scored 4.82, while GPT-o3 scored 4.92 in the same reported comparison. The small gap after quantization is operationally useful, but it is still a result on this corpus and this evaluator.
BLEU supplies the reported overlap comparison. The paper does not state the implementation signature, tokenization, casing, or smoothing configuration, so exact independent reproduction requires information outside the article. The rubric covers accuracy, fluency, coherence, style, and cultural adaptation, but inherits judge bias. A 40-fable human pilot adds a useful check, not a population-level validation: it had one evaluator.
Scope of the result
In one inspected example, the untuned model changed a skunk into an invented creature; the tuned model preserved the animal. That error is more revealing than a fourth decimal place of BLEU. Literary translation can remain fluent while silently changing what happened.
The release keeps four records distinct: the 15K training set, the three-million-pair corpus, the adapter configuration, and the evaluation run. That separation is what makes the 4.43-to-4.83 comparison auditable.