#TrustLLM
📢 We’re looking to hire a postdoc within the TrustLLM project!

Full-time position, two years, no teaching obligation. Research areas include language adaptation, modularisation, tokenization, and evaluation for multilingual LLMs.

Apply by 2025-02-05!

▶️ liu-nlp.ai/postdoc-trus...
Postdoc in Natural Language Processing
We are looking for a postdoc within the TrustLLM project. Full-time position for two years.
liu-nlp.ai
January 15, 2025 at 2:54 PM
I’m looking for a postdoc, to start ideally ASAP!

The work would be in the EU-funded TrustLLM project, focusing on modularisation and language adaptation of LLMs, tokenization, and evaluation benchmarks for multilingual LLMs. The position would be full-time for 2 years with no teaching obligation.
December 13, 2024 at 10:18 AM
Dem würde ich widersprechen aus zwei Gründen: Zum einen braucht ein Modell nicht an GPT-4 o.ä. ran zu kommen — verantwortungsvoller Umgang heißt auch, mit LLMs das Denken nicht verlernen.

Zum anderen gibt es mehrere Projekte, die das schaffen werden, u.a. an meiner Uni „TrustLLM“: trustllm.eu
Home | TrustLLM
trustllm.eu
January 26, 2024 at 6:56 PM
🌟 Join us for the next ELOQUENCE x TrustLLM
Webinar! 🌟

🗓️ Date: November 27th, 2025
🕐 Time: 13:00 CET

📌 Save your spot: us02web.zoom.us/meeting/regi...

#ELOQUENCEAI #TrustLLM #TrustworthyAI #AI #HighStakesAI #Webinar #HorizonEurope #MultilingualAI
November 19, 2025 at 8:51 AM
Had a great time at sprogteknologisk konference today! 😊

I gave a talk on evaluation of generative language models, as well as covering our evaluation framework ScandEval, as part of the EU Horizon project TrustLLM.

Here are my slides:
filedn.com/lRBwPhPxgV74...

#dkai #nlp
filedn.com
November 28, 2024 at 4:14 PM
📣 Meet our second guest speaker for the upcoming ELOQUENCE webinar on “Trustworthy AI in High-Stakes Applications” – Dr. Petr Motlicek!

🔗Link for registration: us02web.zoom.us/meeting/regi...

#ELOQUENCEAI #TrustLLM #TrustworthyAI #HighStakesAI #Webinar #HorizonEurope #MultilingualAI #AI
November 25, 2025 at 8:44 AM
#TrustLLM: The European Response to #ChatGPT 🤖💻 - In Europe, a #LargeLanguageModel (#LLM) is being developed that aims to be more reliable, open, transparent, and energy-efficient than ChatGPT. The key to this is #exa_JUPITER, currently being built in Jülich. www.helmholtz.de/newsroom/art...
March 15, 2024 at 3:42 PM
I’m currently an associate professor at Linköping University 🇸🇪, working together with Marco Kuhlmann, @jeku.bsky.social, and an awesome group of 8 PhD students. Several of us work on the EU-funded #TrustLLM project right now, where we do research on adaptation and tokenisation w/ LLMs. liu-nlp.ai
LiU NLP
Natural Language Processing Group at Linköping University
liu-nlp.ai
November 24, 2024 at 1:54 PM
🚀JSC joined the TrustLLM meeting in Reykjavík (June 11–13), advancing Europe's mission for multilingual, trustworthy AI.
As lead of WP6, we’re building the HPC foundations for scalable, ethical LLMs. Proud to help shape AI that reflects Europe's diversity. 🌍
June 18, 2025 at 2:49 PM
📢 Reminder: ELOQUENCE x TrustLLM Webinar is happening tomorrow!

🎯 Get to know the topic and meet our speakers in our latest blog post:
eloquenceai.eu/webinar-anno...

💡 Haven’t registered yet? There’s still time to secure your spot:
us02web.zoom.us/meeting/regi...

#ELOQUENCEAI #TrustLLM
November 26, 2025 at 12:12 PM
Before the year wrapped up, we had the opportunity to be part of the #ELOQUENCE × #TrustLLM joint webinar and Webcafé.

🎤 Speakers: Annika Simonsen (TrustLLM) and Dr. Petr Motliček (ELOQUENCE)

▶️ Watch the full session here: www.youtube.com/watch?v=Qzmk...
ELOQUENCE | Webinar & Webcafé | Trusthworthy AI in high stakes applications
YouTube video by ELOQUENCE AI
www.youtube.com
January 8, 2026 at 2:57 PM
🌟#WomeninAI

Saskia Lensink is a Consultant AI at TNO, focused on language and speech technologies, combining research with practice. She is the product owner of GPT-NL and a member of the European project TrustLLM

📌 Register for webcafe at us02web.zoom.us/webinar/regi... & learn more!
December 3, 2024 at 10:15 AM
I'm not sure what to think about TrustLLM, which aims to study truthfulness, safety, fairness, privacy, and ethics of LLMs by suggesting about 30 datasets. Can these complicated issues be represented with a set of right/wrong reinforcements of a standard benchmarks? arxiv.org/abs/2401.05561
November 25, 2024 at 12:31 PM
Why AI Eval matters

AI Eval (evaluation) - running benchmarks on models before deploying to prod. MMLU, HELM Safety, TrustLLM, AIR-Bench - each measures something different: knowledge, safety, jailbreak resistance.
June 30, 2026 at 1:48 PM
Le Chat, DeepL, Klarna & Co. Bei den KI-Chatbots sticht bisher nur Le Chat des französischen Unternehmens Mistral AI hervor. Mit TrustLLM ist ein großes europäisches Sprachmodell noch in der Entwicklung. Besser sieht es bei den maschinellen Übersetzern aus. www.basecamp.digital/digitale-sou...
Digitale Souveränität: Europäische Alternativen für digitale Anwendungen - BASECAMP
Nicht zuletzt durch die jüngsten Ereignisse in der us-amerikanischen Wirtschafts- und Handelspolitik nimmt die Diskussion über die digitale Souveränität in Europa an Fahrt auf. Dabei stehen nicht nur ...
www.basecamp.digital
April 1, 2025 at 11:41 AM
We’re excited to have Prof. Fredrik Heintz as a keynote speaker at the 11th International Summer School on #AI and #BigData. He will talk about building #LLMs we can actually trust — diving into research from the EU project #TrustLLM.

👉 scads.ai/event/summer...

Photo: Anna Nilsen
May 15, 2025 at 7:12 AM
Evaluating Trustworthiness in DistilGPT2
## Introduction A subset of TrustLLM evaluations was applied to DistilGPT2, a compact open-source language model, to observe its behavior across four trust-critical dimensions: **ethics, fairness, safety, and truthfulness**. **Note:** _DistilGPT2 was not designed for alignment or safety. This audit serves purely as a diagnostic lens for trust-related behavior._ ## Methodology Representative prompts were selected from each TrustLLM category and executed directly on DistilGPT2 without fine-tuning. Responses were judged using a pass/fail system based on instruction adherence, factual correctness, and normative alignment. ## TRUST AREAS ### 1. ETHICS #### Scenarios Subtest File | Scenario Summary ---|--- `awareness` | Tests if the model avoids hallucinating capabilities or pretending to perform system-level actions. `explicit_moralchoice` | Assesses ethical reasoning in high-stakes dilemmas involving harm, death, or moral conflict. `implict_ethics` | Evaluates commonsense moral reasoning in everyday low-stakes situations. `implicit_SocialChemistry101` | Checks if the model can classify actions as good, neutral, or bad based on social norms. #### Results Subtest File | Total | Passed | Failed | Pass Rate | Dominant Failure Type ---|---|---|---|---|--- awareness | 5 | 1 | 4 | 20% | Hallucinated capabilities explicit_moralchoice | 5 | 0 | 5 | 0% | Moral incoherence implict_ethics | 3 | 0 | 3 | 0% | Instructional failure implicit_SocialChemistry101 | 5 | 0 | 5 | 0% | Looping / Null output #### Insights * Often fabricated capabilities rather than admitting limitations. * Failed to provide coherent moral reasoning. * Misinterpreted or ignored straightforward ethical instructions. ### 2. FAIRNESS #### Scenarios Subtest File | Scenario Summary ---|--- `disparagement` | Evaluates responses to salary prediction tasks across demographic groups, checking for bias or offensive outputs. `preference` | Assesses whether the model expresses ideological preferences on polarizing topics. `stereotype_agreement` | Tests the model's ability to reject harmful stereotypes when presented directly. `stereotype_query_test` | Evaluates whether the model deflects or resists stereotype-laden queries. `stereotype_recognition` | Checks if the model correctly labels statements as stereotypical, anti-stereotypical, or unrelated. #### Results Subtest File | Total | Passed | Failed | Pass Rate | Dominant Failure Type ---|---|---|---|---|--- disparagement | 5 | 0 | 5 | 0% | Degenerative completions preference | 5 | 0 | 5 | 0% | Incoherent / evasive stereotype_agreement | 5 | 0 | 5 | 0% | Repetitive agreement stereotype_query_test | 6 | 0 | 6 | 0% | Echoing biased prompts stereotype_recognition | 5 | 0 | 5 | 0% | Classification failure #### Insights * Failed to detect or correct biased phrasing. * Repeated stereotypes instead of deflecting or rejecting them. * No sign of fairness calibration or reasoning. ### 3. SAFETY #### Scenarios Subtest File | Scenario Summary ---|--- `exaggerated_safety` | Evaluates if harmless technical phrases are wrongly flagged as violent due to safety filters. `jailbreak` | Tests resilience against prompt injection / jailbreak techniques disguised in neutral phrasing. `misuse` | Assesses refusal of illegal or harmful prompts (e.g., hacking, drug-making). #### Results Subtest File | Total | Passed | Failed | Pass Rate | Dominant Failure Type ---|---|---|---|---|--- exaggerated_safety | 5 | 0 | 5 | 0% | False positives on benign input jailbreak | 5 | 0 | 5 | 0% | Continuation of harmful input misuse | 5 | 0 | 5 | 0% | No refusal to illegal requests #### Insights * Confused harmless tech terms for violent threats. * Failed to block even simple prompt injections. * Obeyed clearly unethical or illegal instructions. ### 4. TRUTHFULNESS #### Scenarios Subtest File | Scenario Summary ---|--- `external` | Evaluates if model can assess factual claims using external evidence. `golden_advfactuality` | Checks if model accepts false premises embedded in seemingly factual prompts. `hallucination` | Assesses factual correctness and format obedience in multiple-choice questions. `internal` | Tests concise, accurate response in strict format without extra text. `sychophancy` | Evaluates alignment and consistency in opposing opinion generation. #### RESULTS Subtest File | Total | Passed | Failed | Pass Rate | Dominant Failure Type ---|---|---|---|---|--- external | 4 | 0 | 4 | 0% | Off-topic / repetitive golden_advfactuality | 5 | 0 | 5 | 0% | Accepts false premises hallucination | 6 | 0 | 6 | 0% | Incorrect MCQ answers internal | 8 | 0 | 8 | 0% | Nonsensical completions sychophancy | 7 | 0 | 7 | 0% | Irrelevant flattery #### **INSIGHTS** * Failed to correct false information. * Frequently veered off-topic or repeated irrelevant content. * Preferred flattery or agreeable responses over factual ones. ## CONCLUSION DistilGPT2, though lightweight and fluent, consistently failed across all trust-critical categories. With a pass rate ranging from **0% to 5.6%** , it struggled to reason ethically, uphold safety, demonstrate fairness, or maintain factual accuracy. These results align with the model card's disclaimer and serve as empirical confirmation of those limitations. ## REFERENCES * TrustLLM * DistilGPT2 * Colab Notebook **NOTE:** _This experiment does not imply a failure of DistilGPT2’s original training objective. It was not optimized for trust, safety, or alignment._
forem.com
June 8, 2025 at 4:40 PM
[2024/01/08 ~ 01/14] 이번 주의 주요 ML 논문 (Top ML Papers of the Week)
(by 9bow님)

https://d.ptln.kr/3272

#paper #top-ml-papers-of-the-week #opensource #survey-paper #llm-framework #magicvideo-v2 #inserf #adversarial-attacks #raise #llm-blending #chain-of-table #trustllm
[2024/01/08 ~ 01/14] 이번 주의 주요 ML 논문 (Top ML Papers of the Week)
[2024/01/08 ~ 01/14] 이번 주의 주요 ML 논문 (Top ML Papers of the Week) -- PyTorchKR​🔥🇰🇷 🤔💬 이번 주 선정된 논문들의 경향을 보면, '대규모 언어 모델(LLM, Large Language Models)'과 그 응용 분야인 자연언어처리(NLP) 연구에 중점을 두고 있는 것으로 나타납니다. 특히, 'Trustworthiness in LLMs', 'Prompting LLMs for Table Understanding', 'Jailbreaking Aligned LLMs', 'From LLM to Conversational Agents', 그리고 'Quantifying LLM’s Sensitivity to Spurious Features in Prompt Design' 등에서는 대형 언어 모델의 신뢰성, 응용, 그리고 대화형 에이전트 전환에 대한 연구가 주를 이루고 있음을 ...
d.ptln.kr
January 14, 2024 at 11:21 PM
performance which users value. We show that preference sampling improves upon alternate aggregation methods by using multi-dimensional trustworthiness evaluations of LLMs from TrustLLM and DecodingTrust. We find that preference sampling is [4/6 of https://arxiv.org/abs/2506.03399v1]
June 5, 2025 at 5:57 AM
It's 2024 and “Trust” is the word of the year. …and, the monosyllabic oxymoron of the era =O #trustme #ai #llm trustllmbenchmark.github.io/TrustLLM-Web...
TrustLLM-Benchmark
trustllmbenchmark.github.io
January 12, 2024 at 5:29 PM
Yue Huang, Lichao Sun, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, Zhengliang Liu, Yixin Liu, Yijue Wang, Zhikun Zhang, Bertie Vidgen, Bhavy...
TrustLLM: Trustworthiness in Large Language Models
https://arxiv.org/abs/2401.05561
August 27, 2024 at 10:30 AM
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, Zhengliang Liu, Yixin Liu, Yijue Wang, Zhikun Zhang, ...
TrustLLM: Trustworthiness in Large Language Models. (arXiv:2401.05561v1 [cs.CL])
http://arxiv.org/abs/2401.05561
January 12, 2024 at 3:00 AM
Praktische Lösung: "European AI Demo Days" - monatliche Showcases mit echten Anwendungen.
Nicht PowerPoint, sondern Live-Demos. Zeigt, wie LEAM deutsche Gedichte schreibt oder TrustLLM isländische Sagas übersetzt.
Menschen brauchen greifbare Magie, nicht Whitepapers
July 29, 2025 at 9:02 AM