1. blocktype: evaluation
2. MCTS(?)
3. "Break this problem into parts."
4. RetroInstruct being named after backtranslation, though the cognitive operation isn't actually taught in the dataset yet(?)
1. blocktype: evaluation
2. MCTS(?)
3. "Break this problem into parts."
4. RetroInstruct being named after backtranslation, though the cognitive operation isn't actually taught in the dataset yet(?)
minihf.com/posts/2024-0...
minihf.com/posts/2024-0...
x.com/cyjustinchen...
x.com/cyjustinchen...
We mused on whether human derived COR biosignatures could be used in preclinical models - turns out - it depends! #TBSky
pubmed.ncbi.nlm.nih.gov/39651886/
We mused on whether human derived COR biosignatures could be used in preclinical models - turns out - it depends! #TBSky
pubmed.ncbi.nlm.nih.gov/39651886/
• Featured a “quote” from “Canadian cryptozoologist Dr. Zach Antel“ (this person appears to be made up, I cannot find a single mention of him anywhere else. Unless there was an JPN -> ENG backtranslation error with his name?) saying that the creature
• Featured a “quote” from “Canadian cryptozoologist Dr. Zach Antel“ (this person appears to be made up, I cannot find a single mention of him anywhere else. Unless there was an JPN -> ENG backtranslation error with his name?) saying that the creature
#AugMix #AutoAugment #BackTranslation #CutMix #Cutout #DataAugmentation #EDA
doyouknow.kr/631/data-aug...
#AugMix #AutoAugment #BackTranslation #CutMix #Cutout #DataAugmentation #EDA
doyouknow.kr/631/data-aug...
多くの大規模言語モデル(LLM)は、有害なリクエストを拒否するように訓練されているが、有害な意図を隠すために元のプロンプトを書き換えるジェイルブレイク攻撃にはまだ脆弱である。この論文では、``backtranslation''によって脱獄攻撃からLLMを防御する新しい方法を提案する。具体的には、入力プロンプト...
多くの大規模言語モデル(LLM)は、有害なリクエストを拒否するように訓練されているが、有害な意図を隠すために元のプロンプトを書き換えるジェイルブレイク攻撃にはまだ脆弱である。この論文では、``backtranslation''によって脱獄攻撃からLLMを防御する新しい方法を提案する。具体的には、入力プロンプト...
2402.16459, cs․CL | cs․AI, 26 Feb 2024
🆕Defending LLMs against Jailbreaking Attacks via Backtranslation
Yihan Wang, Zhouxing Shi, Andrew Bai, Cho-Jui Hsieh
2402.16459, cs․CL | cs․AI, 26 Feb 2024
🆕Defending LLMs against Jailbreaking Attacks via Backtranslation
Yihan Wang, Zhouxing Shi, Andrew Bai, Cho-Jui Hsieh