Project 05 of 13

Clinical Research Assistant — RAG & Dataset Assessment

Retrieval is only part of the application.

  • Current client work
  • Application & data services
Read the case study

The problem

A clinical research project with two connected areas: a research application that answers from study documents with source context, and separate assessment services for data-labelling work.

What I did

  • Section-based chunking, embeddings, vector retrieval, and semantic plus keyword search.
  • LLM integration, streamed answers, cancellation handling and report exports.
  • Data-labelling assessments: test authoring, access rules, media questions and response review.
  • S3/R2 storage and an Electron agent for local dataset indexing, validation and media access.

Decisions

  1. Chunk by section

    Study context stays with each retrieved piece.

  2. Search both ways

    Semantic and keyword search together, so paraphrases and exact terms both match.

  3. Let people stop the answer

    Streaming comes with cancellation.

  4. Assess datasets, not models

    Assessments support dataset review.

Boundaries

Client project under my current Upwork contract: backend services and selected frontend workflows, built with AI coding agents. Related services, not one runtime. Assessment is not clinical validation or a measured benchmark of answer quality.

Synthetic demo · fictional data