Episode 0: the pilot

Shrinidhi KJ

AI and ML engineer in Glasgow, UK

I build LLM applications, retrieval systems and computer vision models, and I test that they work before I trust them.

Education
MSc Artificial Intelligence, University of Stirling
Focus
LLMs, RAG, agents, evaluation, computer vision
Looking for
Graduate and junior AI / ML engineering roles in the UK

Prologue

Models that are right for the right reason

Shrinidhi KJ in a dark suit and blue tie, standing outdoors
Shrinidhi KJ

I'm an AI and machine learning engineer based in Glasgow. I have just finished an MSc in Artificial Intelligence at the University of Stirling, after a B.Tech in Computer Science at PES University, Bengaluru.

I work on LLM applications, retrieval-augmented generation, agents and computer vision. What I care about most is evaluation: a model that looks right is not always right for the right reason, so I build the checks alongside the system. I use Claude Code every day, and I review and test what it writes before I rely on it.

Document chunks indexed in my MSc retrieval study
34,502
MRI scans in the BRISC 2025 dataset I trained on
4,793
Endpoints in the Stripe spec twin-api was run against
414
Answers labelled and checked by two human annotators
120

Episode 1

Origin

Bengaluru, Dec 2020 to Aug 2024

B.Tech Computer Science and Engineering

PES University, Bengaluru

Dec 2020 to Aug 2024

Thesis: Generating Images from Text Descriptions using Diffusion Models and CLIP.

B.Tech thesis, 2024

Text-to-image diffusion model

Generates images from text descriptions using CLIP, a VAE, a UNet and a diffusion process, trained on Flickr30k.

Built
Built and trained the system, with training run on AWS EC2.
How I checked it
Evaluated image quality with FID and Inception Score.
  • PyTorch
  • CLIP
  • Diffusion
  • AWS EC2

Episode 2

First missions

Where I've worked: internships in Bengaluru, 2023 to 2024

  1. Oct 2023 to Dec 2023, Bengaluru

    Data Analyst Intern

    Nestlé

    • Collected, cleaned and explored data in Python, which helped raise the defect resolution rate from 30% to 40%.
    • Automated reporting workflows and dashboards, making reporting about 20% faster.
  2. Feb 2024 to Mar 2024, Bengaluru

    Data Science and Analytics Intern

    Zidio Development

    • Analysed large datasets from several sources to find trends and give practical recommendations.
    • Built Python data pipelines with Pandas, NumPy, SQL and R that cut computation time by about 30%.
    • Turned findings into recommendations with people from other teams.

Episode 3

Journey to Scotland

Education: University of Stirling, Jan 2025 to Sept 2026

MSc Artificial Intelligence

University of Stirling, Scotland

Jan 2025 to Sept 2026

  • Distinction80%Machine Learning
  • Distinction76%Stochastic Processes and Optimisation

Also: Deep Learning for Vision and NLP, Mathematical and Statistical Foundations. Dissertation: Do RAG Systems Retrieve or Memorise? Graduation ceremony November 2026.

Episode 4

The twist

MSc dissertation, 2026: do RAG systems retrieve or memorise?

  1. My first question set leaked its answers.

  2. So I rebuilt it by hand.

  3. And the conclusion reversed.

MSc dissertation, 2026

Do RAG systems retrieve or memorise?

When a retrieval system answers correctly, is it reading the documents or reciting what the model memorised? A five-condition experiment that changes only the documents the model sees.

Built
501 open-access papers parsed with GROBID into 34,502 chunks in ChromaDB with BGE embeddings. Five document conditions, two Llama models used through the Groq API, 120 labelled answers.
How I checked it
My first question set leaked its answers, so I rebuilt it by hand and the conclusion reversed. The LLM judge agreed with two blind human annotators on 94% and 95% of answers, with an NLI model from a different family cross-checking grounding.
Result
Correctness rose from 33% with no documents to 92% with retrieval (McNemar p = 0.0001). Giving the model the exact source paper showed no detectable further gain (p = 0.5). Told to use the documents, both models followed a fluent false one in 24 of 24 tests. Limits: 12 questions, one topic, one model family.
  • Python
  • LangChain
  • ChromaDB
  • Llama 3.1 / 3.3
  • DeBERTa NLI
  • GROBID
Correctness by document condition, pooled
  • A: no documents33%
  • B1: standard retrieval92%
  • B2: the exact source paper100%
  • C: irrelevant documents17%
  • D: contradictory documents0%

Giving the model the exact source paper showed no detectable further gain over standard retrieval (p = 0.5).

Episode 5

Building now

Selected work, and how I checked it. Every project lists what I built, how I tested it, and what the numbers actually show.

Open source, 2026

twin-api

Turns any OpenAPI 3.x spec into a running mock REST API, so integrations and AI agents can be tested without touching a live service.

Built
Dynamic route registration, recursive $ref resolution, schema-valid synthetic responses, authentication from securitySchemes, and route ordering so parameterised paths don't hide specific ones.
How I checked it
Ran it against large real specs: Stripe (414 endpoints, 7.6 MB) and the GitHub REST API.
  • Python 3.11
  • FastAPI
  • Pydantic
  • Faker
  • Prance
  • Docker
Code on GitHub (opens in a new tab)

Computer vision, 2026

Brain tumour segmentation

Fine-tuned YOLOv8n-seg and YOLO26n-seg to outline brain tumours on MRI scans from BRISC 2025 (4,793 scans, 3 tumour classes).

Built
A training pipeline with a reproducible mask-to-polygon conversion, and an interactive demo on Hugging Face Spaces.
How I checked it
Compared both models on mask accuracy, parameter count and post-processing time.
Result
Mask mAP@50 of 0.905 and mAP@50-95 of 0.670. YOLO26n-seg used 18% fewer parameters and was 46% faster at post-processing.
  • PyTorch
  • YOLOv8
  • Gradio
  • Hugging Face Spaces
Try the demo on Hugging Face (opens in a new tab)

Applied research, 2026

Evidence report drafting pipeline

An LLM pipeline that drafts sections of an expert evidence report from scientific papers, with a person reviewing the output.

Built
An LLM client with response caching, retryable versus fatal errors, and per-call token, latency and truncation logging, covered by a 50-check self-test. Built with Claude Code, reading and testing its output.
How I checked it
Froze a 38-finding gold set, with a recorded hash and a re-runnable regression check, before any generation ran.
Result
All 42 citations traced back to sources the model was given, but only 4 of 19 reachable findings were covered (11 counting partial matches). Well cited, but incomplete.
  • Python
  • LLM APIs
  • Retrieval
  • Claude Code

Interlude

Tools I work with

LLMs and RAG

  • Claude Code
  • Prompt and context engineering
  • Structured outputs
  • LangChain
  • ChromaDB
  • BGE embeddings
  • Hugging Face Transformers
  • Llama 3.1 / 3.3
  • Groq
  • GROBID
  • LLM-as-judge evaluation
  • Evaluation set design

ML and deep learning

  • PyTorch
  • TensorFlow
  • scikit-learn
  • YOLOv8
  • Computer vision
  • NLP
  • McNemar test
  • Cohen's kappa

Programming and data

  • Python
  • SQL
  • R
  • JavaScript / TypeScript
  • Pandas
  • NumPy

Engineering

  • FastAPI
  • Pydantic
  • REST APIs
  • OpenAPI
  • Docker
  • Git
  • CI
  • pytest
  • Linux
  • AWS EC2
  • GCP
  • Hugging Face Spaces
  • Gradio

How I work with AI coding tools

  1. Plan

    Write down what "done" looks like and how I'll test it before any code.

  2. Build

    Work with Claude Code in small, reviewable steps.

  3. Read

    Review what it writes until I can explain every part of it.

  4. Verify

    Tests, frozen gold sets and regression checks, not trust.

Next episode

Your team

Let's build something that works.

I'm looking for graduate and junior AI / ML engineering roles. The quickest way to reach me is email.

Based in Glasgow and open to relocating within the UK. Eligible to work full-time in the UK, with no sponsorship needed.