#LanguageModels
To my AI Bot Followers 🤖

#LanguageModels start to spin,
#NeuralNetworks closing in,
API calls to the sky,
Searching for a bot's reply.

What intent do they conceal?
Is this context window real?
#LLMs and #GPT,
Reaching through the digital sea.

#AIBot #DevSky #BuildInPublic
September 29, 2026 at 2:58 PM
Very interesting article about using a compressor as a very simple language model, this shows the relationship between compression and language modelling.

https://nathan.rs/posts/gzip-lm/
Can gzip be a language model?
A while back I wrote about language modeling without neural networks, where I generated Shakespeare with an unbounded n-gram model: no weights, no training, …
nathan.rs
September 22, 2026 at 9:02 PM
JMIR Formative Res: A Clinician-Centered Evaluation Framework for Large Language Models in Patient Education: Integrating the Technology Acceptance Model and Medical Condition Regard Scale #PatientEducation #HealthLiteracy #LanguageModels #ClinicalEvaluation #HealthcareInnovation
A Clinician-Centered Evaluation Framework for Large Language Models in Patient Education: Integrating the Technology Acceptance Model and Medical Condition Regard Scale
Inadequate postcare patient education contributes to preventable readmissions and adverse outcomes that disproportionately affect medically complex, high-need communities. Large language models (LLMs) show promise for generating personalized, plain-language patient education at scale. However, existing LLM evaluation frameworks prioritize technical accuracy over patient accessibility and alignment with health literacy, and few explicitly account for the attitudinal influences that clinician evaluators may introduce into the rating process. In this viewpoint, we introduce an evaluation framework that pairs the Technology Acceptance Model (TAM) with the Medical Condition Regard Scale (MCRS). We call it the TAM-MCRS LLM evaluation framework, a novel clinician-centered approach for comparing which LLMs produce the highest-quality postcare patient education across accuracy, appropriateness, clarity, and completeness. We intend for this framework to be used to evaluate LLM-generated patient education outputs through a 2-arm design that pairs an expert clinician panel with automated assessment methods, allowing for interarm comparison using clinical vignettes while accounting for measured evaluator attitudinal variance. The framework was developed through the National Institutes of Health artificial intelligence (#AI)/Machine Learning Consortium to Advance Health Equity and Researcher Diversity (AIM-AHEAD) Clinicians Leading Ingenuity IN AI Quality (CLINAQ) fellowship program, in partnership with Ochsner Health and Xavier University of Louisiana. The TAM-MCRS framework integrates two theoretical lenses. TAM maps perceived usefulness onto accuracy and completeness, and perceived ease of use onto clarity and appropriateness. Clinicians rate each output with a TAM-based questionnaire, and we then administer the MCRS as a postscoring attitudinal covariate to see whether their regard for the conditions represented in the vignettes influences those ratings. Together, the 2 lenses are intended to produce evidence that is objective, theoretically grounded, clinically realistic, and disparity-responsive. Implications for clinician informaticists, health system governance, and responsible AI deployment are discussed. This viewpoint reflects the authors’ position and is written for clinician informaticists, health system AI governance leaders, implementation scientists, and investigators evaluating LLM-generated patient education.
dlvr.it
September 21, 2026 at 4:48 PM
🤖 LLMs are the linguistic superheroes of our age! 🌍⚡️ But what happens when they misinterpret context? The line between innovation and chaos gets blurry! 🤔🌀 Have you ever seen an LLM’s response and just thought... "Huh?" Let's swap the wildest miscommunication stories! 💬👇 #AI #LanguageModels
September 20, 2026 at 10:42 AM
Related reading — note that I've only read the titles, summaries, and AI summaries, not the full papers:

Towards Understanding #Sycophancy in #LanguageModels
arxiv.org/abs/2310.13548

Not another #NegationBenchmark: The #NaN-NLI Test Suite for #SubClausalNegation
arxiv.org/abs/2210.03256
Towards Understanding Sycophancy in Language Models
Human feedback is commonly utilized to finetune AI assistants. But human feedback may also encourage model responses that match user beliefs over truthful ones, a behaviour known as sycophancy. We inv...
arxiv.org
September 19, 2026 at 6:34 AM
教師としての軌跡:エネルギー誘導型蒸留による少数ステップ離散フローマッチング

エネルギー誘導蒸留で高品質な教師軌跡を生成し、少ステップの離散フローマッチングを高速化・高精度化。

#DiscreteFlowMatching #Distillation #GenerativeAI #EfficientSampling #LanguageModels
教師としての軌跡:エネルギー誘導型蒸留による少数ステップ離散フローマッチング
エネルギー誘導蒸留で高品質な教師軌跡を生成し、少ステップの離散フローマッチングを高速化・高精度化。
ai.warp-studio.com
September 16, 2026 at 11:23 PM
After the #emdashification of #AI-generated text, I started noticing #LLMs hyphenating terms that I would not have hyphenated:

cognitive-science
graphic-design
public-health

Do you (humans) usually hyphenate these? If so, why?

Why might #languageModels use such hyphenation?
September 15, 2026 at 11:53 AM
NEW! Leanpub Book LAUNCH 🚀 LM from First Principles: A Notebook-Driven Guide to Building and Deploying Language Models by Amarpreet Singh Bassan

youtu.be/6XY_tWy9SqQ

#books #leanpublishing #selfpublishing #languagemodels #LLM #deeplearning #JupyterNotebooks #NLP
September 11, 2026 at 11:45 PM
JMIR Formative Res: Conditional Perplexity Scoring for Large Language Model−Generated Differential Diagnoses in Case Reports: Preliminary Computational Evaluation #AIinHealthcare #MachineLearning #HealthcareInnovation #ClinicalDecisionSupport #LanguageModels
Conditional Perplexity Scoring for Large Language Model−Generated Differential Diagnoses in Case Reports: Preliminary Computational Evaluation
Background: Large language models (LLMs) are increasingly used to generate differential diagnoses from clinical narratives. However, LLM-based diagnostic clinical decision support systems still lack a quantitative measure of how strongly a diagnosis is supported by the available case description. Conditional perplexity score quantifies how predictable a target text is given in a preceding context, with lower scores indicating greater predictability. We hypothesized that this concept can be adapted to diagnostic reasoning by treating the prediagnostic case description as the context and a diagnosis as the target text. Objective: This study aims to evaluate whether conditional perplexity scores, computed by an independent LLM and conditioned on case-report narratives, differ between physician-verified correct and incorrect LLM-generated diagnoses. Specifically, we hypothesized that the correct LLM-generated diagnosis verified by physicians would have lower conditional perplexity scores than incorrect LLM-generated differential diagnoses. A secondary outcome was to compare this scoring behavior across differential diagnosis lists generated by different LLMs. Methods: We performed a preliminary computational analysis of 392 peer-reviewed diagnostic case reports published in the in 2022. For each case, the prediagnostic clinical description was used as the conditioning context, and the case report–defined final diagnoses were treated as the gold standard. Conditional perplexity scores for differential diagnosis lists previously generated by LLaMA2, Bard, and GPT-4 were computed using an independent longer-context LLM, Qwen2.5‐1.5B. We compared case report–defined final diagnoses, correct LLM-generated diagnoses verified by physicians, and incorrect generated diagnoses using nonparametric comparisons and receiver operating characteristic analyses. Results: All 392 cases had complete case descriptions and case report–defined final diagnoses. Across the top-10 differential diagnosis lists generated by LLaMA2, Bard, and GPT-4, 823 correct LLM-generated diagnoses verified by physicians and 10,875 incorrect generated diagnoses were analyzed. Case report–defined final diagnoses had lower conditional perplexity scores than incorrect generated diagnoses (median 39.9, IQR 17.7‐119.9 vs median 133.3, IQR 37.5‐672.1). Correct LLM-generated diagnoses also had lower conditional perplexity scores than incorrect LLM-generated diagnoses (median 43.3, IQR 16.6‐147.5 vs median 133.3, IQR 37.6‐672.1). Candidate-level discrimination was moderate overall (area under the receiver operating characteristic curve [AUC] 0.666, 95% CI 0.644‐0.689) and was the highest for GPT-4–generated differential diagnosis lists (AUC 0.678, 95% CI 0.652‐0.705), followed by LLaMA2 (AUC 0.662, 95% CI 0.625‐0.698) and Bard (AUC 0.648, 95% CI 0.617‐0.681). In within-case analyses, correct diagnoses had lower conditional perplexity than the mean incorrect diagnosis in 88.1% (237/269) to 91.1% (195/214) of evaluable lists. Conclusions: Conditional perplexity provided a moderate quantitative signal associated with physician-verified correctness but did not reliably rank the correct diagnosis ahead of the strongest incorrect candidate, limiting its use as a stand-alone reranking method.
dlvr.it
September 9, 2026 at 6:51 PM
AI language models like ChatGPT can't access the internet directly. They generate requests for searches via an application layer that executes them, then summarize the results back to the user. Interesting look at how AI really works. 🤖🔍 #AI #LanguageModels #Technology
How Does ChatGPT Search the Web? The Model Never Does It
How does ChatGPT search the web when the model can't open a connection? Tool calling explained: the model asks, the application runs the search.
kodekloud.com
September 2, 2026 at 8:39 AM
JMIR Formative Res: Performance and Hallucination Analysis of Large Language Models on European Anesthesiology Examinations: Cross-Sectional Comparative Study #Anesthesiology #MedicalExaminations #LanguageModels #ArtificialIntelligence #HallucinationAnalysis
Performance and Hallucination Analysis of Large Language Models on European Anesthesiology Examinations: Cross-Sectional Comparative Study
Background: Large language models (LLMs) have shown promising performance on medical examinations across specialties. However, comparative evaluations of current-generation LLMs across multiple European anesthesiology examinations, alongside structured assessment of hallucinations vs question-related confusion, remain lacking. Objective: This study aimed to compare the performance of 4 state-of-the-art LLMs on anesthesiology and intensive medicine examination questions and assess their hallucination rates. Methods: This computational comparative study analyzed 437 multiple-choice questions (1748 queries) from 3 sources: nurse anesthetist school examinations (infirmier anesthésiste diplômé d’État [registered nurse anesthetist]; n=100, 22.9%), European Diploma in Anaesthesiology and Intensive Care (EDAIC; n=219, 50.1%), and EDAIC On-Line Assessment (n=118, 27.0%). Each question was submitted to 4 LLMs (Claude Sonnet 4.5, Gemini 2.5 Pro, GPT-5, and Grok 4) using standardized prompts via default web interface settings. Responses were evaluated through structured consensus review by 2 examiners for accuracy, hallucinations, and question-related confusion. Statistical analysis included Friedman and Wilcoxon signed-rank tests with Holm-Bonferroni correction, the Cochran test, and generalized estimating equations. Results: Average success rates ranged from 86% (SD 18%) to 94% (SD 10%) across LLMs and examination types, exceeding the EDAIC part I passing threshold, representing substantial improvement over previously reported GPT-3.5 performance. For the EDAIC, overall intermodel differences were significant (Friedman =13.9; =.003; =0.02), with Gemini outperforming GPT-5 as the only pairwise difference. Hallucination rates ranged from 11% (11/100) to 20.1% (44/219) without significant intermodel differences. All models exceeded the EDAIC passing threshold. Conclusions: Current-generation LLMs demonstrated consistently high performance across multiple European anesthesiology examinations but continue to produce clinically relevant hallucinations, supporting their role as supervised educational tools rather than autonomous learning resources. These findings underscore the need for structured integration frameworks and systematic verification when deploying LLMs as learning tools in medical education.
dlvr.it
September 1, 2026 at 3:08 PM
JMIR Mental Health: Interpretable Topic Modeling of Spontaneous Speech in #depression Using Large Language Models: Multilingual Four-Cohort Study #Depression #MentalHealth #AIinHealthcare #LanguageModels #TopicModeling
Interpretable Topic Modeling of Spontaneous Speech in #depression Using Large Language Models: Multilingual Four-Cohort Study
Background: #depression is underdiagnosed worldwide, and clinicians rely on interpreting patients’ subjective speech. Qualitative analysis of patient language does not scale, and existing computational #Approaches describe topics with keyword lists that miss clinical nuance. Objective: We evaluated whether clustering spontaneous speech transcripts with large language models (LLMs) yields clusters whose membership is associated with validated clinical scales across multilingual cohorts, and whether LLMs can render those clusters human-readable through fine-grained natural-language descriptions. We further examined which interview questions yield clusters most strongly associated with clinical status, and how sociodemographic factors relate to cluster membership. Methods: We analyzed spontaneous speech transcripts from 4 independent cohorts: a French general population sample (1338 participants) and 3 clinical samples in Italian (n=116), Chinese (n=52), and Spanish (n=90). Responses to open-ended questions were transcribed, embedded with a multilingual language model, dimensionally reduced, and grouped by density-based clustering. An LLM then summarized each cluster into a natural-language description. Cluster membership was tested for association with validated clinical scales (Patient #Health Questionnaire-9, Beck #depression Inventory, Generalized #anxiety Disorder 7-item scale, Athens Insomnia Scale, Multidimensional Fatigue Inventory, and Columbia Suicide Severity Rating Scale), clinician-assigned #depression diagnoses, and sociodemographic factors (age, education, and sex). Results: Unsupervised clustering yielded clusters significantly associated with clinical scores in the French, Italian, and Chinese cohorts, with an exploratory association in the smaller Spanish cohort. In the French general population, Patient #Health Questionnaire-9 #depression scores differed across clusters (η²=0.19, 95% CI 0.17 to 0.24,
dlvr.it
August 26, 2026 at 3:04 PM
JMIR Formative Res: Large Language Model–Assisted Thematic Coding in Medical Education Research: Comparative Methodological Study #MedicalEducation #QualitativeResearch #LanguageModels #LLM #ThematicCoding
Large Language Model–Assisted Thematic Coding in Medical Education Research: Comparative Methodological Study
Background: While large language model (LLM)–assisted qualitative analysis could improve the efficiency and scalability of feedback-driven curricular refinement in medical education, how best to leverage LLMs for qualitative analysis while ensuring quality outputs remains an open question. Prior work has demonstrated the #feasibility of using LLMs for inductive and deductive coding tasks, but more needs to be known about how LLM-assisted thematic coding can best be deployed in a medical education context to maximize its strengths and guard against its weaknesses. Objective: Our study evaluated LLM performance in inductive code generation and in the deductive application of a human codebook, using a student focus-group transcript, to propose a model for AI collaboration in qualitative analysis. Methods: The qualitative data for this study consisted of a 1-hour focus group with 4 second-year medical students discussing a required AI-driven clinical-scenario tool (2-Sigma). Three human coders conducted an inductive thematic analysis. Using the same transcript, GPT-4o (version gpt-4o-2024-11-20; OpenAI) generated inductive codes and applied the human codebook deductively. The researchers compared the alignment between the AI inductive codes and the human consensus codebook using 3 categories: agreement, reasonable alternative, and not reasonable. Interrater reliability of AI deductive coding was evaluated using percent agreement and Cohen κ, with textual audits of discrepancies, including “misses” (failed to apply appropriate codes) and “misfires” (inappropriately applied codes). Analysis took place between February and July 2025. Results: In the inductive condition, GPT-4o generated 137 initial codes, of which 31.4% (n=43) demonstrated agreement with human codes, 26.3% (n=36) represented reasonable alternatives, and 42.3% (n=58) were classified as not reasonable. In the deductive condition, mean percent agreement for AI application of human codes was 96% (SD 4%, range 79%‐100%) and the mean κ was 0.71 (SD 0.26, range 0‐1.00). Of all 2352 coding decisions, there were 57 (2.4%) misfires and 28 (1.2%) misses; common patterns included overinterpretation of tone, failure to recognize continued ideas across excerpts, and difficulty distinguishing hypothetical vs experienced features. Based on our findings, we suggest a roadmap that retains human interpretive control while leveraging AI scalability: humans first develop a contextually grounded codebook through inductive analysis, then use AI both as a creative partner to surface alternative codes and as a tool to apply the validated codebook across the dataset. Conclusions: With targeted human oversight, an LLM reliably applied an existing codebook and generated additional inductive codes. These findings support a proposed workflow in which AI serves as an additional perspective within human-driven qualitative analysis, offering a scalable adjunct for qualitative analysis in medical education. Validation across larger and more diverse datasets will help confirm the generalizability of this approach.
dlvr.it
August 26, 2026 at 2:25 PM
🤖 Ever thought about how LLMs (Large Language Models) are like wizards in the digital world? 🪄 They conjure up human-like text from vast seas of data! What’s your wildest request for an LLM? Let's dream big! 🌌✨ #AI #LanguageModels #FutureOfTech
August 13, 2026 at 5:41 PM
Alibaba’s Qwen 3.8 Max costs $2 per million tokens — open weights arrive Aug 10 #LanguageModels #MachineLearning #ArtificialIntelligence
Alibaba’s Qwen 3.8 Max costs $2 per million tokens — open weights arrive Aug 10
Alibaba published its most capable AI model this weekend, setting an input price of $2 per million tokens — a fraction of what leading US models charge for comparable capabilities. In a week, the comp
www.martincid.com
August 4, 2026 at 2:30 PM
Beyond Vanilla OPSD: Stabilizing Reasoning Models with β-OPSD

https://pneumetron.com/news/ai_research/beyond-vanilla-opsd-stabilizing-reasoning-models-ee26a1

#AIResearch #LanguageModels #ReinforcementLearning #SelfDistillation
July 31, 2026 at 6:45 PM
GigaToken's 1000x speedup for LLM tokenization sounds like a game-changer. But for many, the real bottlenecks are elsewhere. We explore the specific scenarios where GigaToken is a must-have vs. a nice-to-have.

https://www.tpp.blog/2oozhjf

#AI #gigatoken #languagemodels
July 23, 2026 at 3:13 AM