Supervised Fine-Tuning or In-Context Learning? Evaluating LLMs for Clinical NER

February 20, 2026

Reading time: 1 minute

...

📝 Original Info

Title: Supervised Fine-Tuning or In-Context Learning? Evaluating LLMs for Clinical NER
ArXiv ID: 2510.22285
Date: 2025-10-25
Authors: John Smith, Jane Doe, Michael Johnson

📝 Abstract

We study clinical Named Entity Recognition (NER) on the CADEC corpus and compare three families of approaches: (i) BERT-style encoders (BERT Base, BioClinicalBERT, RoBERTa-large), (ii) GPT-4o used with few-shot in-context learning (ICL) under simple vs.\ complex prompts, and (iii) GPT-4o with supervised fine-tuning (SFT). All models are evaluated on standard NER metrics over CADEC's five entity types (ADR, Drug, Disease, Symptom, Finding). RoBERTa-large and BioClinicalBERT offer limited improvements over BERT Base, showing the limit of these family of models. Among LLM settings, simple ICL outperforms a longer, instruction-heavy prompt, and SFT achieves the strongest overall performance (F1 $\approx$ 87.1%), albeit with higher cost. We find that the LLM achieve higher accuracy on simplified tasks, restricting classification to two labels.

Supervised Fine-Tuning or In-Context Learning? Evaluating LLMs for Clinical NER

📝 Original Info

📝 Abstract

💡 Deep Analysis

📄 Full Content

Reference

Table of Contents

Table of Contents

📝 Original Info

📝 Abstract

💡 Deep Analysis

📄 Full Content

Reference

Related Posts

A Bi-population Particle Swarm Optimizer for Learning Automata based Slow Intelligent System

Active Learning Techniques for Pomset Recognizers

Assessing a mobile-based deep learning model for plant disease surveillance

Start searching

No results found