CuraLit: Benchmarking and Optimizing LLMs for Biomedical Literature RAG Support to facilitate new research projects

Project Information

ai, artificial-intelligence, big-data, bioinformatics, biology, computer-science, data, python, Rust
Project Status: Recruiting
Project Region: PA Science
Submitted By: Oliver Bonham-Carter
Project Email: obonhamcarter@allegheny.edu
Project Institution: Allegheny College
Anchor Institution: CR-Penn State
Project Address: 520 North Main Street
Meadville, Pennsylvania. 16335

Preferred Start Date: As soon as possible.

Mentors: Oliver Bonham-Carter
Students: Magnolia Myers

Project Description

Project Description:

CuraLit is a Rust-based biomedical literature analysis software tool that helps undergraduate and or novice researchers to find and identify relevant PubMed articles, visualise research trends, and build domain-specific retrieval-augmented generation (RAG) systems for literature-driven discovery. Researchers interact with RAG systems that they create to investigate research opportunities in keeping with the goals of their studies. For its RAG support, Large Language Models (LLMs) are required.

The proposed project is to extend the currently developing CuraLit system by testing and benchmarking it with diverse LLMs and retrieval strategies to improve speed, accuracy, and usability for undergraduate researchers working in research areas of molecular, cellular, and computational biology, and for biomedical studies that require a literature review.

This remote undergraduate research project provides student mentorship and training in biomedical data science, AI evaluation, and reproducible computational research. The student researcher will work with faculty who will provide IT support to complete computational experiments and testing. Some of the goals of this work will be to determine how different LLMs perform when used to create biomedical literature workflows, ascertain working software settings to support practical use in education and research.

Problem Statement:

A common barrier in undergraduate research during the early stages of biological and biomedical research is the inability to efficiently navigate the large collections of literature. For example, when researching a method for testing a hypothesis, article keywords may provide few details about their contents because the article keywords for search engines may have been be poorly chosen, or due to relevant information that may have been omitted from the abstract. This frustration grows when the researcher attempts (and perhaps fails) to synthesize information from diverse sets of articles into ideas for working research questions. Existing AI tools may help with creating literature reviews, however without careful benchmarking such technologies may be slow, inconsistent, or simply unreliable for the discovery of ideas and evidence to begin a structured project. CuraLit addresses this challenge by combining PubMed searching, statistical analyses, visualizations, and local RAG workflows to help the researcher get started. Before it is ready for use, CuraLit will require rigorous testing across multiple LLMs, the identification of appropriate retrieval settings, and a testable validation that the outputs are useful and appropriate.

While CuraLit is still in its developmental stages, the tool will eventually be released as an open source project. The proposed project will evaluate model choice, retrieval design, citation reliability, and determine software settings that serve to strengthen CuraLit as a useful resource for researchers during the early stages of their biomedical projects.

Collaborators:

- Oliver Bonham-Carter, PhD, Allegheny College, Project Lead / Faculty Mentor / IT Technical Mentor
- Undergraduate Student: Magnolia Myers, Allegheny College, Undergraduate Researcher

Research Plan and Expected Outcomes:

The student will evaluate multiple LLMs and retrieval configurations using CuraLit’s existing PubMed workflow, including local model deployment and RAG pipelines. The study will compare models based on answer quality, retrieval relevance, citation accuracy, latency, and reproducibility across biomedical research questions. The student will report settings in the software that help to facilitate project creation in biological and biomedical research.

The project will pursue three aims:

1. Evaluate different LLMs and retrieval settings for biomedical question answering using PubMed-derived evidence.
2. Optimize the CuraLit RAG workflow for speed, citation reliability, and user accessibility.
3. Develop a reproducible benchmark and documentation framework for future student researchers.

Expected outcomes include a comparative model benchmark, a recommended production workflow, improved RAG performance, and a reusable evaluation framework and associated settings to boost support for literature reviews in biomedical projects. The student will also gain direct experience in literature analysis, computational methods, and AI model evaluation in a real research setting.

Timeline:

Project duration: 3 to 6 months, depending on student availability

- Stage 1: Environment setup, baseline evaluation, and model selection criteria.
- Stage 2: LLM/RAG benchmarking, retrieval tuning, and performance testing.
- Stage 3: Final validation, documentation, reporting, and presentation.

Deliverables:

- Comparative benchmarking across multiple LLMs and retrieval configurations.
- Recommended settings for a fast, accurate, and reproducible CuraLit workflow.
- Reusable scripts and documentation for future model evaluation and benchmarking.
- A report of "lessons learned" to help guide CuraLit users to get started.
- An open source and well documented CuraLit project
- Final project report and presentation to the project team and the PA Science DMZ/NCEMS community.

Significance:

This project is motivated to advance the mission of NRRE-P2 by integrating biomedical research, data science, and mentoring in a remote collaborative environment. By improving CuraLit’s LLM and RAG evaluation process, the project will help undergraduate students develop more efficient and evidence-based research workflows in biomedical sciences and bioinformatics.

Additional Resources

Github Contributions: https://github.com/developmentAC/curalit
Wrap Presentation: 3-4 months

Project Information

ai, artificial-intelligence, big-data, bioinformatics, biology, computer-science, data, python, Rust
Project Status: Recruiting
Project Region: PA Science
Submitted By: Oliver Bonham-Carter
Project Email: obonhamcarter@allegheny.edu
Project Institution: Allegheny College
Anchor Institution: CR-Penn State
Project Address: 520 North Main Street
Meadville, Pennsylvania. 16335

Preferred Start Date: As soon as possible.

Mentors: Oliver Bonham-Carter
Students: Magnolia Myers

Project Description

Project Description:

CuraLit is a Rust-based biomedical literature analysis software tool that helps undergraduate and or novice researchers to find and identify relevant PubMed articles, visualise research trends, and build domain-specific retrieval-augmented generation (RAG) systems for literature-driven discovery. Researchers interact with RAG systems that they create to investigate research opportunities in keeping with the goals of their studies. For its RAG support, Large Language Models (LLMs) are required.

The proposed project is to extend the currently developing CuraLit system by testing and benchmarking it with diverse LLMs and retrieval strategies to improve speed, accuracy, and usability for undergraduate researchers working in research areas of molecular, cellular, and computational biology, and for biomedical studies that require a literature review.

This remote undergraduate research project provides student mentorship and training in biomedical data science, AI evaluation, and reproducible computational research. The student researcher will work with faculty who will provide IT support to complete computational experiments and testing. Some of the goals of this work will be to determine how different LLMs perform when used to create biomedical literature workflows, ascertain working software settings to support practical use in education and research.

Problem Statement:

A common barrier in undergraduate research during the early stages of biological and biomedical research is the inability to efficiently navigate the large collections of literature. For example, when researching a method for testing a hypothesis, article keywords may provide few details about their contents because the article keywords for search engines may have been be poorly chosen, or due to relevant information that may have been omitted from the abstract. This frustration grows when the researcher attempts (and perhaps fails) to synthesize information from diverse sets of articles into ideas for working research questions. Existing AI tools may help with creating literature reviews, however without careful benchmarking such technologies may be slow, inconsistent, or simply unreliable for the discovery of ideas and evidence to begin a structured project. CuraLit addresses this challenge by combining PubMed searching, statistical analyses, visualizations, and local RAG workflows to help the researcher get started. Before it is ready for use, CuraLit will require rigorous testing across multiple LLMs, the identification of appropriate retrieval settings, and a testable validation that the outputs are useful and appropriate.

While CuraLit is still in its developmental stages, the tool will eventually be released as an open source project. The proposed project will evaluate model choice, retrieval design, citation reliability, and determine software settings that serve to strengthen CuraLit as a useful resource for researchers during the early stages of their biomedical projects.

Collaborators:

- Oliver Bonham-Carter, PhD, Allegheny College, Project Lead / Faculty Mentor / IT Technical Mentor
- Undergraduate Student: Magnolia Myers, Allegheny College, Undergraduate Researcher

Research Plan and Expected Outcomes:

The student will evaluate multiple LLMs and retrieval configurations using CuraLit’s existing PubMed workflow, including local model deployment and RAG pipelines. The study will compare models based on answer quality, retrieval relevance, citation accuracy, latency, and reproducibility across biomedical research questions. The student will report settings in the software that help to facilitate project creation in biological and biomedical research.

The project will pursue three aims:

1. Evaluate different LLMs and retrieval settings for biomedical question answering using PubMed-derived evidence.
2. Optimize the CuraLit RAG workflow for speed, citation reliability, and user accessibility.
3. Develop a reproducible benchmark and documentation framework for future student researchers.

Expected outcomes include a comparative model benchmark, a recommended production workflow, improved RAG performance, and a reusable evaluation framework and associated settings to boost support for literature reviews in biomedical projects. The student will also gain direct experience in literature analysis, computational methods, and AI model evaluation in a real research setting.

Timeline:

Project duration: 3 to 6 months, depending on student availability

- Stage 1: Environment setup, baseline evaluation, and model selection criteria.
- Stage 2: LLM/RAG benchmarking, retrieval tuning, and performance testing.
- Stage 3: Final validation, documentation, reporting, and presentation.

Deliverables:

- Comparative benchmarking across multiple LLMs and retrieval configurations.
- Recommended settings for a fast, accurate, and reproducible CuraLit workflow.
- Reusable scripts and documentation for future model evaluation and benchmarking.
- A report of "lessons learned" to help guide CuraLit users to get started.
- An open source and well documented CuraLit project
- Final project report and presentation to the project team and the PA Science DMZ/NCEMS community.

Significance:

This project is motivated to advance the mission of NRRE-P2 by integrating biomedical research, data science, and mentoring in a remote collaborative environment. By improving CuraLit’s LLM and RAG evaluation process, the project will help undergraduate students develop more efficient and evidence-based research workflows in biomedical sciences and bioinformatics.

Additional Resources

Github Contributions: https://github.com/developmentAC/curalit
Wrap Presentation: 3-4 months