Back to research
CSE443: Bioinformatics · First author · 2025

BioAlign-QLoRA: Biomedical Knowledge Graph Alignment

Does fine-tuning a general LLM on biomedical facts actually reorganise how it represents biology? QLoRA-tuned Llama-3, Mistral, and Phi-3 on 68,444 gene–disease associations, then measured both accuracy and the geometry of their embeddings.

First author on a team coursework project · Concept, data curation, all code, fine-tuning, and the separation metric

[ EVD ]

Engineering evidence

My contribution
First author · Data curation, QLoRA experiments, evaluation design
Outcome
Custom embedding-alignment metric and a three-model comparative study
Decision record 01
Used QLoRA so three model families could be compared within student compute, keeping adapters small enough to publish with the repository.
[ RES ]

Key results

Mistral 7B accuracy
83.8%
Zero-shot gene–disease classification after QLoRA
Llama-3 8B accuracy
81.0%
Zero-shot, after QLoRA
Phi-3 Mini KG separation
+126%
Change in the Knowledge Graph Separation score after tuning
Associations
68,444
Curated from the Comparative Toxicogenomics Database
[ IMG ]

Screenshots · 3 captures

BioAlign-QLoRA: Biomedical Knowledge Graph Alignment screenshot 101 / 03
[ ARC ]

Architecture

The Knowledge Graph Separation score asks whether gene and disease embeddings that are linked in the knowledge graph sit closer together than unlinked ones, a geometric check that accuracy alone can't give.

  1. 01

    CTD extraction

    • Gene–disease associations
    • Evidence counts, classes
  2. 02

    Curation

    • 68,444 cleaned pairs
    • EDA on classes and diseases
  3. 03

    QLoRA fine-tuning

    • 4-bit base models
    • r = 16, α = 32, 2,000 steps
    • Llama-3 · Mistral · Phi-3
  4. 04

    Evaluation

    • Zero-shot accuracy
    • KG Separation score
    • vs BioMistral-7B
[ KEY ]

Technical highlights

  1. 01

    Introduced the Knowledge Graph Separation score to quantify how well an LLM's embedding space matches biomedical knowledge structure.

  2. 02

    Fine-tuned three model families on consumer-grade hardware with QLoRA and compared them against the domain-pretrained BioMistral-7B.

[ LIM ]

Limits & lessons

  • Mistral's accuracy rose while its separation score fell (−38%): accuracy and embedding geometry can disagree, which is exactly why both were measured.
[ STK ]

Stack

  • Python
  • PyTorch
  • Transformers
  • PEFT / QLoRA
  • Unsloth
  • Llama-3
  • Mistral-7B
  • Phi-3