We present a benchmark of 57+ competitive text-based games to evaluate&train LLMs
including negotiation, deception, theory of mind...
Multiplayer support
Human-vs-models
Model-vs-model
Perfect for social interaction, Multi-Agent, multi-turn reasoning and Planning
🤖📈
We present a benchmark of 57+ competitive text-based games to evaluate&train LLMs
including negotiation, deception, theory of mind...
Multiplayer support
Human-vs-models
Model-vs-model
Perfect for social interaction, Multi-Agent, multi-turn reasoning and Planning
🤖📈
We report our findings in our paper "Skill Issue: Are Skills Language-Invariant in LLMs?".
And with it, TextArena is now multilingual 🌍 with 65 games in 193 languages. 🧵
We report our findings in our paper "Skill Issue: Are Skills Language-Invariant in LLMs?".
And with it, TextArena is now multilingual 🌍 with 65 games in 193 languages. 🧵
A blogpost on human-model interaction, games, training and testing LLMs
research.ibm.com/blog/LLM-soc...
🤖📈🧠
A blogpost on human-model interaction, games, training and testing LLMs
research.ibm.com/blog/LLM-soc...
🤖📈🧠
alphaXiv: https://alphaxiv.org/abs/2608.25832
HF Paper: https://huggingface.co/papers/2608.25832
Code: https://github.com/TextArena/TextArena
Here I focus on an analysis of 2D board games checking where exactly do the language performance gaps originate🔍
Here I focus on an analysis of 2D board games checking where exactly do the language performance gaps originate🔍
TextArena
https://arxiv.org/abs/2504.11442
TextArena
https://arxiv.org/abs/2504.11442
Talk to us:
discord.gg/KMndsqwMaZ
git:
github.com/LeonGuertler...
Talk to us:
discord.gg/KMndsqwMaZ
git:
github.com/LeonGuertler...
arxiv.org/abs/2504.11442
More analysis and fuller results soon (this is more about the framework itself, the next more traditional conference submission)
arxiv.org/abs/2504.11442
More analysis and fuller results soon (this is more about the framework itself, the next more traditional conference submission)
"TextArena - A library for training and evaluation of language models in competitive text-based environments [(games)]"
textarena.ai
Each player can get its own language (e.g. player 0 in
Each player can get its own language (e.g. player 0 in
An open-source suite of 57+ games designed to test LLM behaviors: negotiation, deception, and more.
Basically, it’s an AI testing ground for agentic skills.
Read more: lnkd.in/gajaR6FC
An open-source suite of 57+ games designed to test LLM behaviors: negotiation, deception, and more.
Basically, it’s an AI testing ground for agentic skills.
Read more: lnkd.in/gajaR6FC
BabyLM will be a workshop and another pretraining competition, do join!
bsky.app/profile/lcho...
BabyLM will be a workshop and another pretraining competition, do join!
bsky.app/profile/lcho...
Specifically, we adopt a similar framework of stackable wrappers and an environment registry
Here is code example of how you can let GPT-4o-mini and Calude-3.5-haiku compete against each other offline
Specifically, we adopt a similar framework of stackable wrappers and an environment registry
Here is code example of how you can let GPT-4o-mini and Calude-3.5-haiku compete against each other offline
all of those decrease after instruction tuning and RLHF
all of those decrease after instruction tuning and RLHF
Stay tuned
Stay tuned
@babyLMchallenge
offers an interactivity track and a workshop submission space
textarena.ai offers a platform to play with models, test models offline or online with others, or invent games.
@babyLMchallenge
offers an interactivity track and a workshop submission space
textarena.ai offers a platform to play with models, test models offline or online with others, or invent games.