#BackTranslation
A significant challenge was to recover the diversity of unformal French through generalized backtranslation and incorporate the RATP’s internal analytical specialized frameworks inside the model's reasoning traces.
June 18, 2026 at 6:00 PM
Ideal use case might be generalized backtranslation/memorization. Otherwise I see the incentives for RL.
December 27, 2025 at 6:29 PM
RetroInstruct and weave-agent looking very good in this context.

1. blocktype: evaluation
2. MCTS(?)
3. "Break this problem into parts."
4. RetroInstruct being named after backtranslation, though the cognitive operation isn't actually taught in the dataset yet(?)
2/13 We identify 4 key cognitive behaviors that enable successful learning: Verification (checking work), Backtracking (trying new approaches), Subgoal Setting (breaking problems down) & Backward Chaining (working backwards from a goal). Qwen naturally exhibits these, while Llama mostly lacks them.
March 4, 2025 at 6:41 PM
Basically because in practice synthetic data is made through methods like backtranslation that are lossless or it's selected through various quality filters and classifiers to improve faster than the random sampling degrades it.
minihf.com/posts/2024-0...
The RetroInstruct Guide To Synthetic Text Data
minihf.com
November 7, 2025 at 6:19 PM
Ah good. It’s topical but: backtranslation techniques could actually help here.
May 10, 2025 at 8:46 PM
Yeah, now I do backtranslation on purpose sometimes, I just don't let myself think I'm inventing the first-ever wheel! 😂
June 12, 2024 at 8:36 PM
"X is (usually) present on Y condition so we can train inverse X from Y" is a general deep net training tactic, and one of my favorites. Up there with "Start with the correct answers and generate the questions (backtranslation), then flip the order and train to learn to answer those questions."
February 6, 2026 at 7:01 AM
Turns out that you can apply a backtranslation-like technique to improve reasoning in LLMs:
x.com/cyjustinchen...
December 2, 2024 at 8:13 PM
Prompt tip: You can start with a few shot list of examples in a category, ask the model for more examples, then add the best examples to your prompt list and reroll again to get a stronger condition than you started with. The quality of the list should climb then fall as creativity is exhausted.
February 6, 2026 at 7:45 AM
Par ailleurs, la backtranslation n'est jamais un bon mécanisme pour juger de la qualité d'une traduction. Source : mon baccalauréat et mes années d'expérience en traduction.
April 3, 2025 at 8:07 PM
This is a mystery. Consistent pastoral theme but they did Heinäpellonpuisto (Hayfield park) into Hejnåkersparken, which… means nothing. Hej, nåker!? Would be Höåkersparken or Höängsparken. It is possible, that the explanation is a local backtranslation of the Finnish name, but weird nonetheless.
November 8, 2024 at 9:25 AM
Defending LLMs against Jailbreaking Attacks via Backtranslation https://arxiv.org/abs/2402.16459
March 25, 2024 at 10:31 AM
Solving Prompt Injection via Backtranslation Discussion
Defending LLMs against Jailbreaking Attacks via Backtranslation
Although many large language models (LLMs) have been trained to refuse harmful requests, they are still vulnerable to jailbreaking attacks, which rewrite the original prompt to conceal its harmful intent. In this paper, we propose a new method for defending LLMs against jailbreaking attacks by ``backtranslation''. Specifically, given an initial response generated by the target LLM from an input prompt, our backtranslation prompts a language model to infer an input prompt that can lead to the response. The inferred prompt is called the backtranslated prompt which tends to reveal the actual intent of the original prompt, since it is generated based on the LLM's response and is not directly manipulated by the attacker. We then run the target LLM again on the backtranslated prompt, and we refuse the original prompt if the model refuses the backtranslated prompt. We explain that the proposed defense provides several benefits on its effectiveness and efficiency. We empirically demonstrate that our defense significantly outperforms the baselines, in the cases that are hard for the baselines, and our defense also has little impact on the generation quality for benign input prompts.
arxiv.org
February 27, 2024 at 1:40 PM
I guess Russian doesn't have "tarnation," I phrased it as "what in tarnation" and apparently the literal backtranslation of the russian phrase there is "what the hell is this?"
July 7, 2026 at 8:42 PM
Well, I'm a horny person and I want to try out the idea the day after tomorrow. It's just that the backtranslation in the translator seriously distorted the negative part.
July 26, 2026 at 8:22 AM
Backtranslation of human RNA biosignatures of tuberculosis disease risk into the preclinical pipeline is condition dependent https://www.biorxiv.org/content/10.1101/2024.06.21.600067v1
Backtranslation of human RNA biosignatures of tuberculosis disease risk into the preclinical pipeline is condition dependent https://www.biorxiv.org/content/10.1101/2024.06.21.600067v1
It is not clear whether human progression to active tuberculosis disease (TB) risk signatures are vi
www.biorxiv.org
June 22, 2024 at 5:16 AM
The time when AI translation was as bad as that is long behind us. ChatGPT translated this (my backtranslation of your text) as: "Are they completely out of their minds? An AI translation would be about as useful as a chocolate teapot. The whole project would go completely to hell."
August 21, 2026 at 6:08 AM
I’ll temporarily break my usual Klingon-only rule to provide an overly literal backtranslation for the benefit of anyone outside of my usual Klingon-speaking audience who might see this. It’s only the first verse which I did just for fun, if there’s any serious interest I can do the other verses:
November 25, 2024 at 3:59 PM
Really fortunate to have worked with Hannah Painter on this story where she put up with my shenanigans > 3 years + Mat leave.

We mused on whether human derived COR biosignatures could be used in preclinical models - turns out - it depends! #TBSky

pubmed.ncbi.nlm.nih.gov/39651886/
Backtranslation of human RNA biosignatures of tuberculosis disease risk into the preclinical pipeline is condition dependent - PubMed
Understanding the strengths or limitations of back-translating human-derived correlate of risk (COR) RNA signatures into the preclinical pipeline may help streamline down-selection of therapeutic vacc...
pubmed.ncbi.nlm.nih.gov
December 11, 2024 at 8:20 PM
• (The RSPCA guy was horrified by its appearance)
• Featured a “quote” from “Canadian cryptozoologist Dr. Zach Antel“ (this person appears to be made up, I cannot find a single mention of him anywhere else. Unless there was an JPN -> ENG backtranslation error with his name?) saying that the creature
September 10, 2026 at 5:39 AM
The Greek seems to be "Huperoksiaun" but I think it's possible that "Pantoxiana" was a rough backtranslation of "Transoxiana", cf. "pantomime"
August 6, 2026 at 6:32 PM
데이터 증강 완벽 가이드! 과적합 방지와 모델 성능 향상의 핵심 기법. 이미지 증강: Mixup vs CutMix vs AugMix 비교, AutoAugment vs RandAugment 성능/비용. NLP 증강: Back Translation, EDA 4가지 연산, LLM 활용 최신 기법. GAN 기반 증강으로 민감도 10% 향상!

#AugMix #AutoAugment #BackTranslation #CutMix #Cutout #DataAugmentation #EDA
doyouknow.kr/631/data-aug...
데이터 증강 완벽 가이드: AI 학습 데이터가 부족할 때의 마법! Mixup, CutMix, AutoAugment 총정리
데이터 증강 완벽 가이드! 과적합 방지와 모델 성능 향상의 핵심 기법. 이미지 증강: Mixup vs CutMix vs AugMix 비교, AutoAugment vs RandAugment 성능/비용. NLP 증강: Back Translation, EDA 4가지 연산, LLM 활용 최신 기법. GAN 기반 증강으로 민감도 10% 향상!
doyouknow.kr
December 5, 2025 at 12:26 PM
2402.16459
多くの大規模言語モデル(LLM)は、有害なリクエストを拒否するように訓練されているが、有害な意図を隠すために元のプロンプトを書き換えるジェイルブレイク攻撃にはまだ脆弱である。この論文では、``backtranslation''によって脱獄攻撃からLLMを防御する新しい方法を提案する。具体的には、入力プロンプト...
February 29, 2024 at 12:06 AM
[27/30] 64 Likes, 46 Comments, 1 Posts
2402.16459, cs․CL | cs․AI, 26 Feb 2024

🆕Defending LLMs against Jailbreaking Attacks via Backtranslation

Yihan Wang, Zhouxing Shi, Andrew Bai, Cho-Jui Hsieh
February 29, 2024 at 12:05 AM