Tundra Space

Tundra Space

Clinical Research Directory

Browse clinical research sites, groups, and studies.

Back to Studies
RECRUITING
NCT07795645

Validating a Medical AI Fine-Tuning Platform for Major Diseases

Sponsor: Beijing Friendship Hospital

View on ClinicalTrials.gov

Summary

The goal of this observational study is to leverage the abundant patient resources and standardized medical records from Beijing Friendship Hospital, Xuanwu Hospital, and Beijing Anzhen Hospital, combined with the existing data and knowledge platform of guidelines, consensus, medical literature, and dialogue data from Beijing Haitian Ruisheng Science Technology Co.,Ltd, with Beijing Zhilan Medical Technology Co., Ltd. conducting the fine-tuning, optimization, and validation of the medical large language model. The model is fine-tuned according to the consultation and diagnostic needs of different departments to improve the quality and efficiency of hospital medical services, enhance intelligence, and elevate the level of medical care. Through deployment to hospitals at all levels, it aims to achieve standardized services and support graded diagnosis and treatment. The overall research includes medical big data construction, medical knowledge graph construction, medical large model training and fine-tuning, and large model application platform development and deployment.

Official title: Development and Validation of Fine-Tuning Techniques for Large Models in Gastric, Cardiovascular and Cerebrovascular Diseases

Key Details

Gender

All

Age Range

18 Years - Any

Study Type

OBSERVATIONAL

Enrollment

96000

Start Date

2024-11-01

Completion Date

2027-08-31

Last Updated

2026-08-31

Healthy Volunteers

No

Interventions

OTHER

Smart Medical Big Data Dataset Construction

Integrating multimodal medical data (electronic medical records, medical images and reports, laboratory results, genetic information) with medical guidelines, expert consensus, clinical databases, medical literature, encyclopedias, patents, and doctor-patient dialogue data. Datasets are constructed in phases: pre-training general knowledge datasets (guidelines, textbooks, literature, historical records) and fine-tuning disease-specific datasets (doctor-patient dialogues, disease-specific records, treatment plans, follow-up records). Techniques include data cleaning, standardization, transformation, annotation, and augmentation. The platform adopts a human-machine collaboration strategy to reduce data annotation costs, combining medical experts' professional knowledge with AI capabilities.

OTHER

Large-Scale Medical Knowledge Graph Construction

Extracting multimodal knowledge graphs from guidelines, consensus, and literature using the QLora training framework. The process involves medical knowledge modeling (defining entities, relations, attributes), entity recognition (disease names, drug names), relation extraction (disease-symptom, drug-treatment relationships), attribute extraction (incidence, dosage), and knowledge fusion and completion to build a tens-of-millions-level knowledge graph. The construction incorporates multimodal data integration (text, images, audio) through image recognition and speech recognition technologies.

OTHER

Multi-Disease Medical Large Model Fine-Tuning Platform Construction and Fine-Tuning

Developing fine-tuning techniques for different disease areas using the ChatGLM large model, with integration of expert feedback. Pre-training uses large-scale medical datasets with language modeling tasks. Fine-tuning incorporates physician annotation, data augmentation (multi-task learning, adversarial training), and prompt-tuning to generate disease-specific responses. Deployment applies quantization to reduce model size and improve inference speed. A quality evaluation system is developed as a standardized method for measuring model performance, incorporating customized methods tailored to specific diseases and drawing on existing frameworks such as Ragas and CMB.

OTHER

Medical History Collection System Development

Building an intelligent medical history collection system using voice interaction, natural language processing, and large model technologies. Speech recognition and NLP techniques collect patient information through multi-round dialogue with strategies including questioning, paraphrasing, feeling reflection, and self-disclosure. Physicians supervise and validate the model through Reinforcement Learning from Human Feedback to improve collection efficiency. After collection, the model extracts key information (chief complaint, history of present illness, past medical history, personal history, allergy history) and uses Retrieval-Augmented Generation (RAG) combined with guidelines for preliminary triage.

OTHER

Clinical Auxiliary Diagnosis Research

Building an intelligent clinical auxiliary diagnosis system with multi-source data fusion integrating heterogeneous medical records, laboratory reports, and imaging text reports using OCR and natural language processing (NLP). Using chain-of-thought and reasoning capabilities of large language models combined with knowledge graph retrieval to generate evidence-based diagnostic recommendations. Physicians provide feedback through a data annotation interface to fine-tune the system, with customization settings to meet individual physician needs.

OTHER

Medical Record Generation and Quality Control Research

Building a medical record generation system enabling automatic entry of patient information and automatic medical record generation using large language models for entity recognition, text classification, and semantic understanding. The system generates records compliant with writing standards including chief complaint, history of present illness, past medical history, personal history, marital and reproductive history, family history, and auxiliary examinations. A quality control system automatically detects generated records and, when core fields are missing, alerts physicians through highlighting and submission blocking, prompting them to make necessary additions and modifications.

OTHER

Doctor-Patient Consultation Audio Data Collection and Processing Research

Providing a solution for consultation audio data collection and processing to support intelligent history collection, rapid triage, and clinical auxiliary diagnosis. The collected audio is converted to accurate textual content with preliminary analysis to identify key medical information (symptoms, medications, treatment processes). The study addresses speech recognition challenges including specialized terminology, dialects, accents, and elderly speech characteristics, using refined speech processing algorithms and deep learning approaches. Strict data protection measures (encryption, anonymization) are applied throughout collection, transmission, storage, and processing to ensure patient data security.

OTHER

Cross-Center External Validation

Validating the generalizability of the five intelligent diagnosis and treatment systems (gastric cancer, chronic gastritis, GERD, coronary artery disease, stroke) by transferring the fine-tuned systems from the training center to other centers for external cross-validation. Feasibility and accuracy are evaluated by comparing information collected from large model-patient interactions and physician-patient interactions across different center environments.

Locations (1)

Beijing Friendship Hospital, Capital Medical University

Beijing, China