#InformationExtraction
Thrilled to announce that our research group will be running the TRIPLET challenge at #ISWC2025! 🗾
@rtroncy.bsky.social EURECOM @orange.com @iritoulouse.bsky.social IBM
#SemanticWeb #KnowledgeGraphs #InformationExtraction #AIChallenge
🚀 Challenges Acceptance Notifications are Out🎉

We are excited to announce that 7 challenges have been accepted for ISWC 2025!

Details will be available soon on the conference website! iswc2025.semanticweb.org

👉 Did your challenge proposal get accepted? Share it in the comments below.

#ISWC2025
March 13, 2025 at 1:01 PM
Rex Douglass is a free agent. Rex is a data bulldozer and champion of science, plus no bs palantir overhead. Get him before he gets got.
September 26, 2025 at 2:50 PM
JMIR Mental Health: Practical Guide to Large Language Models for Information Extraction in Behavioral #Health Notes: Tutorial #MentalHealth #HealthTech #NLP #InformationExtraction #ClinicalNotes
Practical Guide to Large Language Models for Information Extraction in Behavioral #Health Notes: Tutorial
Background: #MentalHealth clinical notes contain decision-critical information often absent from structured electronic #Health record fields. Large language models (LLMs) can extract clinically relevant signals from narrative text; however, variability in output format, limited reproducibility, and inconsistent evaluation remain barriers to clinical deployment. Despite rapid advances in LLM-based information extraction, clear and reproducible guidance for interdisciplinary clinical teams is limited. Objective: This tutorial aims to present a structured workflow for zero-shot information extraction from #MentalHealth clinical notes using locally deployed open-source LLMs. It aims to reduce barriers for clinicians and researchers with limited familiarity with natural language processing (NLP) or LLM-based pipelines. Each stage includes key decision points and examples. The workflow is illustrated on two tasks using synthetic notes: (1) detection of self-injurious thoughts and behaviors (SITB) in pediatric emergency department (ED) notes and (2) antipsychotic medication nonadherence detection in outpatient notes, using schema-constrained outputs and standardized evaluation. Methods: We describe a five-stage zero-shot LLM pipeline: (1) infrastructure setup with local deployment via to prevent protected #Health information (PHI) transmission; (2) task definition specifying the clinical construct, output format, and evaluation; (3) dataset preparation using synthetic notes; (4) iterative prompt development using a hold-out development set with binary and Likert scale outputs constrained via JSON schemas; and (5) output parsing, normalization, and validation. We generated 300 synthetic notes per task using separate LLMs for generation and evaluation; 200 notes were used for evaluation, and 100 notes (50 positive and 50 negative) were used as a prompt-development set and excluded from final metrics. Evaluation used Large Language Model Meta AI (Llama) 3.2 and Llama 3.3 with deterministic decoding (temperature=0). Performance was assessed using accuracy, precision, recall, and -score; Likert thresholds were optimized using the Youden index with bootstr#Apped CIs. Results: We demonstrated the pipeline’s functionality using 2 example behavioral #Health detection tasks. Across both examples, the more capable model (Llama 3.3) performed better than the lighter model used earlier in development (Llama 3.2), and we described how the pipeline’s evaluation and error-analysis steps work in practice. These examples also illustrated 2 useful design choices: requiring the model to output in a fixed format reduced errors, and using a graded rating scale, rather than a simple yes/no format, allowed the detection threshold to be adjusted based on clinical risk tolerance. These results are meant to show that the pipeline works as intended, not to serve as a benchmark of real-world accuracy. Conclusions: A schema-driven, zero-shot LLM workflow can support reproducible extraction of clinically relevant information from narrative notes. Local deployment enables processing without transmitting PHI to external servers. This tutorial provides a transferable methodology for institutional adaptation and validation prior to clinical use. All prompts, code, and datasets are publicly available via Zenodo (European Organization for Nuclear Research [CERN]).
dlvr.it
September 24, 2026 at 8:35 PM
Using Information Extraction to Normalize the Training Data for Automatic Radiology Report Generation

PDF 👉 buff.ly/4g92Z3D
I appreciate the authors sharing their #CXRGraph code 👉 buff.ly/3B5V08B

#InformationExtraction #radiology #RadiologyReporting
November 28, 2024 at 6:54 PM
We currently have two fully-funded open PhD positions in our group with a focus on #NLProc, #InformationExtraction and #TextGeneration. I can really recommend both the group as well as Philipp Cimiano as a supervisor, so take this opportunity!
March 6, 2025 at 2:40 PM
The presentation "Extracting Citation Data Using LLMs" by @anwagnerdreas.hcommons.social.ap.brid.gy‬, David Carreto Fidalgo & me on how to extract structured reference information from footnote-heavy scholarship using LLMs:
www.youtube.com/watch?v=bgps...
#LLM #InformationExtraction #Footnotes
Extracting Citation Data Using LLMs | C. Boulanger, D. Carreto Fidalgo & A. Wagner
YouTube video by Network Epistemology in Practice (NEPI)
www.youtube.com
June 11, 2025 at 9:15 AM
What’s your experience with building or scaling KGs in production? Would love to hear your take!
#KnowledgeGraph #RAG #InformationExtraction #Neo4j #AI #NLP
February 13, 2025 at 8:14 AM