C-MAT: Cross-Modal Aligned Transformer
A four-class model that separates Alzheimer's disease, frontotemporal dementia, Parkinson's disease, and healthy controls using structural MRI and resting-state EEG together, and keeps working when a patient has only one of the two.
Team thesis with Sharika Salim, Tonmoy Pal, and Rubayet Hassan Rupom · Supervised by Dr. Chowdhury Mofizur Rahman and Dr. Md. Golam Rabiul Alam
Engineering evidence
- My contribution
- Thesis co-author · Architecture, data pipeline, training, and evaluation
- Outcome
- Defended May 2026 · Grade A · IEEE-format paper prepared
- Decision record 01
- Rendered EEG features as images so a single pretrained vision backbone could serve both modalities, rather than training a separate EEG network on a few hundred recordings.
Key results
- Macro F1 · ConvNeXt-Tiny
- 0.611
- ± 0.029 · subject-level stratified 5-fold CV, four classes
- Paired MRI + EEG subset
- 0.702
- Macro F1 on 25 strictly paired subjects, vs 0.482 for late fusion
- Subjects
- 1,102
- 10 open datasets from seven countries
- Focal vs cross-entropy
- +0.048
- Macro F1 gain from class-balanced focal loss (fold-0 ablation)
Architecture
Both modalities are turned into images so one pretrained vision backbone can read them. Each patient's slices or EEG epochs are pooled with learned attention, and a sigmoid gate decides per patient how much to trust MRI versus EEG. A learnable null token stands in for whichever modality is missing.
01
Cohort & splits
Ten heterogeneous open datasets merged into one audited manifest.
- Dataset-prefixed subject IDs
- Subject-level stratified 5-fold
- No subject in train and test
02
MRI pipeline
T1-weighted volumes normalised across scanners.
- 1 mm resampling, Otsu brain mask
- 14 slices: 7 axial, 4 coronal, 3 sagittal
- CLAHE → 3 × 224 × 224
03
EEG pipeline
Resting-state recordings mapped to 19 standard channels.
- 256 Hz, 1–45 Hz, ICA + ICLabel
- 2 s epochs → 38 features per channel
- Pseudo-image and tabular paths
04
Shared encoder
One pretrained timm backbone for both modalities.
- ConvNeXt-Tiny or DeiT
- Attention pooling over slices / epochs
- EEG feature MLP (19 × 38)
05
Gated fusion
Sigmoid gate mixes the two embeddings per patient.
- Learnable null token per modality
- Class-balanced focal loss + MixUp
- AD · FTD · PD · HC
06
Explainability
- Grad-CAM on MRI slices
- Fusion gate values by class
Technical highlights
- 01
Audited the pipeline through three versions (V1 → V3), fixing subject-ID collisions between datasets, preprocessing reuse, and baseline fairness before trusting any number.
- 02
Handled missing modalities with learnable null tokens instead of imputation, so subjects with only MRI or only EEG still train the shared model.
- 03
Made fusion inspectable: the gate value shows how much the model relied on MRI versus EEG for each class.
- 04
Compared against honest baselines: EEG power spectrum + SVM (0.309) and EEG random forest (0.453).
Limits & lessons
- Research and education only. The model is not a clinical diagnostic tool.
- A strong MRI-only baseline stays competitive on single folds; fusion's clearest gain is on strictly paired subjects (0.702 vs 0.482), which are few.
- The thesis text and figures are © BRAC University, so this page describes the method in my own words instead of reproducing them.
Stack
- Python
- PyTorch
- timm
- ConvNeXt
- DeiT
- MNE
- nibabel
- scikit-learn
- Optuna