Repository for final project allocation and submission for MM60024 : Biomedical Imaging Informatics offered in Spring 2026 at IIT Kharagpur taught by Prof Debashree Guha Adhya.
This repository might be updated with new projects and/or changes to existing projects. Please check back regularly.
Final projects are approved by Prof Debashree Guha Adhya.
- Final Projects for MM60024 : Biomedical Imaging Informatics
- Instructions
- Project Allocation
- Projects
- Project 1 : Classification of Breast Cancer Using Ultrasound Images
- Project 2 : Maternal Health Risk Prediction in Personalized Patient Care
- Project 3 : Early Detection of Coronary Heart Disease
- Project 4 : Cervical Cancer Diagnosis Based on Machine Learning Techniques
- Project 5 : Prostate Cancer Diagnosis Using Machine Learning Methods
- Project 6 : Pneumonia Detection Using Chest X-ray Imaging
- Project 7 : Automatic Tumor Segmentation from Breast Ultrasound Images
- Project 8 : Explainable Survival Prediction of Liver Cirrhosis Patients
- Project 9 : Multiorgan Segmentation Using the AbdomenAtlas Mini Dataset
- Project 10 : Automatic Skin Lesion Segmentation from Dermoscopic Images
- Project 11 : Multimodal Image Fusion of MRI and SPECT Brain Images
- Resources
- All the best!
- The project is to be done in groups of 5 students. The students are expected to work together collaboratively.
- The choice of programming language is left to the students. However, the most common language used is Python.
- Each group will be assigned a mentor TA who will be responsible for guiding the group throughout the project.
- Meetings with the mentor TA will be scheduled at the beginning of the project and at regular intervals.
- Each student will be evaluated based on the contribution towards the project. Make sure you are contributing equally to the project.
- Code plagiarism will not be tolerated. Any submission found to be plagiarized will be awarded a zero grade.
- Late submissions will not be accepted.
- The final project evaluation is based on the following criteria:
Code Quality and Documentation : 60%Final Submission and Report : 40%
Code Quality and Documentation: 60%- This will be based on the following criteria:
- Code Quality : 30% (based on the code quality and readability)
- Documentation : 30% (based on the documentation of the code and the project)
- This will be based on the following criteria:
Final Submission and Report: 40%- This will be based on the following criteria:
- Final Submission : 20% (based on the final submission of the project)
- Final Report : 20% (based on the final report of the project)
- This will be based on the following criteria:
Forkthisgithub.com/SanchitaMondal/BII_MM60024_Projectsrepository.Clonethe forked repository to your local machine using the following command:git clone github.com/{your_username}/BII_MM60024_Projects- Your projects are in the
submissionsdirectory. You can find the project description in the README.md file of the respective project directory. - Work on the project and make
regular commitsto your local repository andpushthem to your forked repository. - Your mentor TA will review your code and provide feedback.
- You have to submit the following:
Final Code: The final code of your project in the respective project directory.- Code should be highly readable and well documented.
- Try to write efficient code and avoid unnecessary code.
Final Report: The final report of your project in the respective project directory. The report should be in the form of amarkdownfile with the namereport.md. The report should contain the following:Introduction: A brief introduction of the project.Data: A brief description of the data used in the project.Questions & Answers: The questions and their respective answers. Also include the code snippets used to answer the questions andwho solvedthe question.References: The references used in the project.
- Submission of the final project will be done via
GitHub Pull Requests. - Once you are done with the project, you can create a
Pull Requestto themainbranch of thegithub.com/SanchitaMondal/BII_MM60024_Projectsrepository. - We will review your merge request and provide feedback. You can make changes to your code and update the merge request. If accepted, your project will be merged to the
mainbranch of thegithub.com/SanchitaMondal/BII_MM60024_Projectsrepository. - That's it!
Congratulations!!have successfully submitted your final project.
- Otherway, you can send a mail with subject
MM60024 : Biomedical Imaging Informatics Final Project (Project number) Submission - Group (Group number)consisting yourFinal Code(.py or .ipynb) andFinal Report(.pdf)
The dates for the final project submission are 15th & 16th April and 30th April 2026, 23:59 IST. (Tentatively)
- This project aims to develop an automated breast cancer classification system using deep learning techniques to distinguish between benign and malignant tumors from BUS images. The dataset is located in the
data/Breast datadirectory. - The study will utilize the BUS-BRA Breast Ultrasound Dataset, which contains annotated ultrasound images with pathology-confirmed labels and lesion segmentation masks.
- Preprocessing techniques : image
resizing,normalization, and dataaugmentationwill be applied to improve data quality and enhance model generalization. - A Convolutional Neural Network (
CNN) or hybrid deep learning architecture will be used to automatically extract important features from the ultrasound images. - The trained model will classify the breast lesions into benign or malignant categories based on the learned features.
- The performance of the proposed model will be evaluated using standard classification metrics such as accuracy, precision, recall, F1-score, and AUC.
- The developed system aims to assist radiologists by providing a computer-aided diagnosis (CAD) tool that improves diagnostic accuracy and supports early detection of breast cancer.
- Data Loading and Initial Exploration: The dataset is located in the
data/Pregnancy_riskdirectory. Load the Excel file containing the maternal health dataset and conduct an initial exploration to understand its structure and contents. - Understand the dataset structure, data types, and basic statistics for each feature.
- The dataset may include clinical features such as:- Age, Systolic Blood Pressure as SystolicBP, Diastolic BP as DiastolicBP, Blood Sugar as BS, Body Temperature as BodyTemp, HeartRate and RiskLevel. All these are the responsible and significant risk factors for high-risk pregnancies.
- The project is divided into following parts:-
Data Preprocessing and Cleaning: A pre-processed dataset with no missing values or inconsistencies, ready for further analysis.Data Analysisis including visualization (e.g., histograms, scatter plots, box plots) and calculating correlations between features.Prediction Task: Predicting Maternal Health Risk Level: Learn classification algorithms such as logistic regression, decision trees, random forests, or support vector machines, and evaluate model performance.- Build a classification model to predict the risk level (low, medium, or high) based on the health features provided.
- Model Tuning and Optimization : Optimize the classification model by tuning hyperparameters and using techniques such as cross-validation.
- For report writing, compile the analysis, and results.
- The objective of this project is to develop a machine learning-based system for early prediction of CHD using patient clinical and lifestyle data. The dataset is located in the
data/Heart_diseasedirectory. - The dataset contains attributes such as
age,gender,blood pressure,cholesterol level,smoking habits,glucose level, and other risk factors. - The project is divided into following parts:-
Data Preprocessing and Cleaning: handling missing values, feature selection, and normalization will be applied to improve data quality and model performance.Data Analysisis including visualization (e.g., histograms, scatter plots, box plots) and calculating correlations between features.Prediction Task: Predict the chances of heart disease in patients: Learn a classification algorithm such as logistic regression, and evaluate model performance using metrics such asaccuracy,precision,recall,F1-score, andconfusion matrix.
- The system can assist healthcare professionals by providing
early warnings for patients at risk of coronary heart disease. - The proposed model aims to support preventive healthcare and clinical decision-making, helping reduce mortality through timely diagnosis and treatment.
- The aim of this project is to develop a machine learning model for predicting the risk of cervical cancer based on patients' demographic, behavioral, and medical history data. The dataset is located in the
data/Cervical datadirectory. - It contains medical and lifestyle information collected from 858 patients at Hospital Universitario de Caracas in Venezuela.
- The dataset contains 36 attributes, including factors such as
age,number of sexual partners,smoking habits,pregnancies,contraceptive use, andsexually transmitted diseases (STDs). - Several diagnostic test results, such as
Hinselmann,Schiller,cytology, andbiopsyare included as target variables to determine the presence of cervical cancer. - The project is divided into following parts:-
Data Preprocessing and Cleaning: handling missing values, feature selection, and normalization will be applied to improve data quality and model performance.Data Analysisis including visualization (e.g., histograms, scatter plots, box plots) and calculating correlations between features.Prediction Task: Detection of cervical cancer: Learn classification algorithms such as Logistic Regression, Decision Tree, Random Forest, or Support Vector Machine will be used to build the prediction model, and evaluate model performance using metrics such asaccuracy,precision,recall,F1-score, andconfusion matrix.
- The developed system can assist healthcare professionals in identifying women at high risk of cervical cancer, enabling early screening and preventive treatment.
- The objective of this project is to develop a machine learning model for predicting prostate cancer based on tumor characteristics extracted from medical data. The data is available at
data/Prostate data - It contains data from 100 patients with multiple tumor-related features and includes 10 attributes, such as
radius,texture,perimeter,area,smoothness,compactness,symmetry, andfractal dimension, which describe tumor characteristics. - The target variable (
diagnosis_result) indicates whether the tumor isbenign(non-cancerous) ormalignant(cancerous). - The project is divided into following parts:-
Data Preprocessing and Cleaning: handling missing values, feature selection, and normalization will be applied to improve data quality and model performance.Data Analysisis including visualization (e.g., histograms, scatter plots, box plots) and calculating correlations between features.Prediction Task: Detection of cervical cancer: Learn classification algorithms such as Logistic Regression, Decision Tree, Random Forest, or Support Vector Machine will be used to build the prediction model, and evaluate model performance using metrics such asaccuracy,precision,recall,F1-score, andconfusion matrix.
- The developed system can assist healthcare professionals by supporting early diagnosis of prostate cancer, which may help improve clinical decision-making and patient care.
- Pneumonia is a serious lung infection that can cause inflammation in the air sacs of the lungs and may become life-threatening if not detected early. The objective of this project is to develop a deep learning-based system to automatically detect pneumonia from chest X-ray images.
- The project uses the Chest X-Ray Pneumonia Dataset, available at
data/Chest X-ray data, which contains labeled chest radiography images for pneumonia diagnosis. - The dataset consists of approximately 5,800 chest X-ray images, divided into two classes: Normal and Pneumonia. The images are organized into training, validation, and test sets, allowing proper training and evaluation of machine learning models.
- Image preprocessing techniques such as
resizing,normalization, anddata augmentationwill be applied to improve model performance and reduce overfitting. - A Convolutional Neural Network (
CNN) will be used to automatically extract features from chest X-ray images. The trained model will classify the input X-ray image into pneumonia-infected or normal lung condition. - Model performance will be evaluated using
accuracy,precision,recall,F1-score, andconfusion matrix. - The developed system can assist doctors and radiologists by providing a computer-aided diagnostic tool for early detection of pneumonia, improving healthcare efficiency and patient outcomes.
- The project's aim is to develop an automated tumor segmentation method that can accurately identify and delineate tumor regions in breast ultrasound images.
- The project uses the BUS-BRA Breast Ultrasound Dataset (
data/Breast data), which contains 60 patients with annotated tumor boundaries and pathology information. - The dataset includes biopsy-confirmed benign and malignant tumors, along with expert-annotated segmentation masks that define the exact tumor region in each image.
- Image preprocessing techniques such as noise reduction, normalization, resizing, and data augmentation will be applied to improve model robustness.
- A deep learning segmentation model, such as
U-Netwill be trained to automatically detect and segment tumor regions in ultrasound images. - The model will learn to classify each pixel as tumor or background, generating a segmentation mask that outlines the tumor boundaries.
- The performance of the segmentation model will be evaluated using
Dice coefficient, andIntersection over Union (IoU). - The developed method can assist radiologists by providing accurate tumor localization and boundary detection, which may support diagnosis, treatment planning, and computer-aided diagnosis systems.
- Liver cirrhosis is a chronic and progressive liver disease that can lead to severe complications and high mortality if not diagnosed and treated in time.
- The objective of this project is to develop an explainable machine learning model to predict the survival outcomes of liver cirrhosis patients using clinical and laboratory data.
- The dataset is available at
data/cirrhosis, which contains patient medical records related to liver cirrhosis and includes multiple clinical features such asage,sex,bilirubin,cholesterol,albumin,platelet count,prothrombin time, anddisease stage, which are important indicators of liver health. - The project is divided into the following parts:-
Data Preprocessing and Cleaning: handling missing values, feature selection, and normalization will be applied to improve data quality and model performance.Data Analysisincludes visualization (e.g., histograms, scatter plots, box plots) and calculating correlations between features.Prediction Task: Learn a machine learning algorithm such as Survival Support Vector Machine will be used to predict patient survival outcomes, and evaluate the prediction model.Explainability: To improve model transparency, Explainable Artificial Intelligence (XAI) techniques such as SHAP (SHapley Additive exPlanations) or LIME (Local Interpretable Model-agnostic Explanations) will be applied to identify the most influential clinical factors affecting survival predictions.
- The proposed explainable system aims to assist clinicians in understanding important risk factors and improving clinical decision-making for liver cirrhosis patient management.
- The objective of this project is to develop a deep learning model for segmenting abdominal organs from CT images to assist medical image analysis and clinical diagnosis.
- The project uses the
AbdomenAtlas1.0Minidataset available on Hugging Face, which contains a large collection of annotated abdominal CT scans. - The dataset includes about 5,195 annotated 3D CT volumes, making it one of the largest publicly available datasets for abdominal organ segmentation.
- Each CT scan contains expert-annotated masks for several abdominal organs such as
liver,spleen,kidneys,stomach,pancreas,gallbladder,aorta, andinferior vena cava (IVC). - Preprocessing steps such as CT slice
normalizationandresizingwill be applied to improve model generalization and training efficiency. - A 3D deep learning segmentation model, such as 3D U-Net will be used to automatically segment multiple abdominal organs from CT volumes. The model will learn voxel-level classification, where each voxel in the CT image is assigned to a specific organ or background.
- The model performance will be evaluated using
Dice Similarity Coefficient (DSC),Intersection over Union (IoU),precision, andrecall, which are commonly used metrics for medical image segmentation. - The proposed system can assist clinicians by automatically identifying abdominal organs in CT scans, supporting tasks such as disease detection, treatment planning, and surgical navigation.
- The aim of this project is to develop a machine learning (deep learning) model for automatic skin lesion segmentation from dermoscopic images, enabling accurate identification of lesion boundaries for early diagnosis of skin cancer.
- The dataset is located in the
data/SkinLesionDatadirectory. It consists of dermoscopic images for training and testing with its ground truth binary segmentation mask, indicating the exact region of the lesion. - The project is divided into the following parts:
Data Preprocessing and Cleaning: Image resizing, preprocessing, normalization, augmentation (such as rotation, flipping), and artifact removal (e.g., hair removal) may be applied to improve model generalization and performance. (if you feel its needed)Data Analysis: Visualization techniques such as sample image inspection, mask overlays, and pixel intensity distributions will be used to understand the dataset. Statistical analysis of lesion sizes and shapes may also be performed.Segmentation Task: Learning models such as U-Net, Fully Convolutional Networks (FCN), Mask R-CNN, or DeepLab may be used for pixel-wise prediction.- Model performance will be evaluated using segmentation-specific metrics such as:
Dice CoefficientJaccard Index (IoU)Pixel AccuracyPrecision and Recall
- The developed system can assist dermatologists by providing precise lesion boundaries, which are crucial for accurate diagnosis and treatment planning.
Here, you are expected to develop a machine learning (or deep learning) based system for multimodal image fusion of MRI and SPECT brain images, combining structural and functional information into a single enhanced image to support improved diagnosis of neurological disorders.
- The dataset is located in the
data/MultimodalFusiondirectory and consists of paired Magnetic Resonance Imaging (MRI) and Single Photon Emission Computed Tomography (SPECT) brain images collected from publicly available medical imaging repositories or clinical datasets. Each pair corresponds to the same patient and registered. - The project is divided into the following parts:
Data Preprocessing and Cleaning: Image registration (alignment of MRI and SPECT images), resizing, normalization, noise reduction, and intensity scaling will be applied to ensure compatibility between modalities and improve fusion quality.Data Analysis: Visualization techniques such as modality comparison, intensity histograms, and overlay analysis will be used to understand differences between MRI and SPECT images. Statistical measures like entropy and mutual information may also be analyzed.Fusion Task: The objective is to generate a fused image that preserves:Deep learning approaches: Convolutional Neural Networks (CNNs), Autoencoders, and Generative Adversarial Networks (GANs)
- The performance of the fusion model will be evaluated using metrics such as:
Entropy (EN)Mutual Information (MI)Structural Similarity Index (SSIM)Peak Signal-to-Noise Ratio (PSNR)
- The developed system can assist healthcare professionals by providing a comprehensive fused view of brain structure and function, which improves the detection and analysis of neurological disorders.
- Python Documentation
- Class Materials
- Python Libraries