Clinical Research Directory
Browse clinical research sites, groups, and studies.
Validating a Medical AI Fine-Tuning Platform for Major Diseases
Sponsor: Beijing Friendship Hospital
Summary
The goal of this observational study is to leverage the abundant patient resources and standardized medical records from Beijing Friendship Hospital, Xuanwu Hospital, and Beijing Anzhen Hospital, combined with the existing data and knowledge platform of guidelines, consensus, medical literature, and dialogue data from Beijing Haitian Ruisheng Science Technology Co.,Ltd, with Beijing Zhilan Medical Technology Co., Ltd. conducting the fine-tuning, optimization, and validation of the medical large language model. The model is fine-tuned according to the consultation and diagnostic needs of different departments to improve the quality and efficiency of hospital medical services, enhance intelligence, and elevate the level of medical care. Through deployment to hospitals at all levels, it aims to achieve standardized services and support graded diagnosis and treatment. The overall research includes medical big data construction, medical knowledge graph construction, medical large model training and fine-tuning, and large model application platform development and deployment.
Official title: Development and Validation of Fine-Tuning Techniques for Large Models in Gastric, Cardiovascular and Cerebrovascular Diseases
Key Details
Gender
All
Age Range
18 Years - Any
Study Type
OBSERVATIONAL
Enrollment
96000
Start Date
2024-11-01
Completion Date
2027-08-31
Last Updated
2026-08-31
Healthy Volunteers
No
Interventions
Smart Medical Big Data Dataset Construction
Integrating multimodal medical data (electronic medical records, medical images and reports, laboratory results, genetic information) with medical guidelines, expert consensus, clinical databases, medical literature, encyclopedias, patents, and doctor-patient dialogue data. Datasets are constructed in phases: pre-training general knowledge datasets (guidelines, textbooks, literature, historical records) and fine-tuning disease-specific datasets (doctor-patient dialogues, disease-specific records, treatment plans, follow-up records). Techniques include data cleaning, standardization, transformation, annotation, and augmentation. The platform adopts a human-machine collaboration strategy to reduce data annotation costs, combining medical experts' professional knowledge with AI capabilities.
Large-Scale Medical Knowledge Graph Construction
Extracting multimodal knowledge graphs from guidelines, consensus, and literature using the QLora training framework. The process involves medical knowledge modeling (defining entities, relations, attributes), entity recognition (disease names, drug names), relation extraction (disease-symptom, drug-treatment relationships), attribute extraction (incidence, dosage), and knowledge fusion and completion to build a tens-of-millions-level knowledge graph. The construction incorporates multimodal data integration (text, images, audio) through image recognition and speech recognition technologies.
Multi-Disease Medical Large Model Fine-Tuning Platform Construction and Fine-Tuning
Developing fine-tuning techniques for different disease areas using the ChatGLM large model, with integration of expert feedback. Pre-training uses large-scale medical datasets with language modeling tasks. Fine-tuning incorporates physician annotation, data augmentation (multi-task learning, adversarial training), and prompt-tuning to generate disease-specific responses. Deployment applies quantization to reduce model size and improve inference speed. A quality evaluation system is developed as a standardized method for measuring model performance, incorporating customized methods tailored to specific diseases and drawing on existing frameworks such as Ragas and CMB.
Medical History Collection System Development
Building an intelligent medical history collection system using voice interaction, natural language processing, and large model technologies. Speech recognition and NLP techniques collect patient information through multi-round dialogue with strategies including questioning, paraphrasing, feeling reflection, and self-disclosure. Physicians supervise and validate the model through Reinforcement Learning from Human Feedback to improve collection efficiency. After collection, the model extracts key information (chief complaint, history of present illness, past medical history, personal history, allergy history) and uses Retrieval-Augmented Generation (RAG) combined with guidelines for preliminary triage.
Clinical Auxiliary Diagnosis Research
Building an intelligent clinical auxiliary diagnosis system with multi-source data fusion integrating heterogeneous medical records, laboratory reports, and imaging text reports using OCR and natural language processing (NLP). Using chain-of-thought and reasoning capabilities of large language models combined with knowledge graph retrieval to generate evidence-based diagnostic recommendations. Physicians provide feedback through a data annotation interface to fine-tune the system, with customization settings to meet individual physician needs.
Medical Record Generation and Quality Control Research
Building a medical record generation system enabling automatic entry of patient information and automatic medical record generation using large language models for entity recognition, text classification, and semantic understanding. The system generates records compliant with writing standards including chief complaint, history of present illness, past medical history, personal history, marital and reproductive history, family history, and auxiliary examinations. A quality control system automatically detects generated records and, when core fields are missing, alerts physicians through highlighting and submission blocking, prompting them to make necessary additions and modifications.
Doctor-Patient Consultation Audio Data Collection and Processing Research
Providing a solution for consultation audio data collection and processing to support intelligent history collection, rapid triage, and clinical auxiliary diagnosis. The collected audio is converted to accurate textual content with preliminary analysis to identify key medical information (symptoms, medications, treatment processes). The study addresses speech recognition challenges including specialized terminology, dialects, accents, and elderly speech characteristics, using refined speech processing algorithms and deep learning approaches. Strict data protection measures (encryption, anonymization) are applied throughout collection, transmission, storage, and processing to ensure patient data security.
Cross-Center External Validation
Validating the generalizability of the five intelligent diagnosis and treatment systems (gastric cancer, chronic gastritis, GERD, coronary artery disease, stroke) by transferring the fine-tuned systems from the training center to other centers for external cross-validation. Feasibility and accuracy are evaluated by comparing information collected from large model-patient interactions and physician-patient interactions across different center environments.
Locations (1)
Beijing Friendship Hospital, Capital Medical University
Beijing, China