Open to AI/ML internships and junior roles

I build ML systems that show their receipts: contamination checks, cited agents, and models you can open and try.

Balamurugan P G/AI/ML Engineer/Coimbatore, Tamil Nadu, India

verascan · train/eval contamination

$

0
public repos
0
live apps
0
tests · Bull/Bear Autopilot
0+
accepted · LeetCode Agent
RAI-Axiom (PPO Sim-to-Real)VeraScan on PyPIDenseNet121 & MobileNetV2Grad-CAM ExplainabilityTrain/Eval Contamination Detection13-Gram & Semantic OverlapMulti-Agent Indian IPO AnalysisDeterministic Fact-Checking568 Automated Tests5-Provider LLM FallbackPlaywright Browser WorkersFastAPI + React + SupabaseStreamlit DashboardsGCP & OCI CloudPython PackagingRAI-Axiom (PPO Sim-to-Real)VeraScan on PyPIDenseNet121 & MobileNetV2Grad-CAM ExplainabilityTrain/Eval Contamination Detection13-Gram & Semantic OverlapMulti-Agent Indian IPO AnalysisDeterministic Fact-Checking568 Automated Tests5-Provider LLM FallbackPlaywright Browser WorkersFastAPI + React + SupabaseStreamlit DashboardsGCP & OCI CloudPython Packaging
01CASE STUDYOngoing

RAI-Axiom

Balamurugan P G and Jason Pandian

A portfolio-allocation policy trained only on synthetic markets, then evaluated zero-shot on six real universes with 10-seed confidence intervals. PPO, no real prices during training. Research by Balamurugan P G and Jason Pandian. Ongoing project; does not claim to beat buy-and-hold (tested tie with SPY out of sample and behind SPY on the holdout point estimate).

PPOReinforcement LearningSynthetic MarketsSim-to-RealZero-Shot Evaluation10-Seed CI

PROBLEM STATEMENT

Train an RL portfolio-allocation policy without exposing it to real historical market prices during training, then test zero-shot transfer across real financial universes.

TECHNICAL METHOD

Trained via Proximal Policy Optimization (PPO) exclusively on synthetic market environments, then evaluated zero-shot on 6 real universes across 10-seed confidence intervals. Ongoing research by Balamurugan P G and Jason Pandian.

VERIFIED RESULT

Tested tie with SPY

out of sample · behind SPY on holdout point estimate (10-seed CI)

02CASE STUDY

VeraScan

Python package that detects train/eval data contamination with 4 checks: exact match, 13-gram overlap, fuzzy match and semantic similarity. Published on PyPI: pip install verascan.

PythonPyPIExact match13-gram overlapFuzzy matchSemantic similarity

PROBLEM STATEMENT

If eval examples leak into training data, benchmark scores look better than the model really is.

TECHNICAL METHOD

Checks train/eval pairs at 4 levels: exact match, 13-gram overlap, fuzzy match and semantic similarity.

VERIFIED RESULT

pip install verascan

Published on PyPI · 4 detection modes

03CASE STUDY

Bull/Bear Autopilot

Multi-agent system for Indian IPO analysis. Every fact is cited to a page, then checked by a deterministic fact-checker; results are delivered on Telegram. 568 tests.

Multi-agent LLMPage-cited factsDeterministic fact-checkTelegram bot

PROBLEM STATEMENT

IPO research is only useful if every claim can be traced back to the source document.

TECHNICAL METHOD

Multiple agents draft bull and bear cases, with each fact cited to a page and then verified by a deterministic fact-checker. Results are delivered on Telegram.

VERIFIED RESULT

568

tests

04CASE STUDY

LungScan AI

5-class chest X-ray classifier built on DenseNet121, with Grad-CAM heatmaps. 86–88% test accuracy. Full-stack app (FastAPI backend, React frontend, Supabase) deployed on GCP, live at medrag.in.

DenseNet121Grad-CAMFastAPIReactSupabaseGCP

PROBLEM STATEMENT

Classify chest X-rays into 5 classes and show which regions drove each prediction.

TECHNICAL METHOD

Fine-tuned DenseNet121 with Grad-CAM overlays, served from FastAPI with a React frontend and Supabase, deployed on GCP.

VERIFIED RESULT

86–88%

test accuracy · 5 classes

05CASE STUDY

AI LeetCode Agent

Agent that solves LeetCode problems using a 5-provider LLM fallback chain and 4 Playwright browser workers. 230+ accepted.

LLM agents5-provider fallbackPlaywright

PROBLEM STATEMENT

Solve LeetCode problems end-to-end without stopping when one LLM provider fails or hits a rate limit.

TECHNICAL METHOD

Falls back across 5 LLM providers, with 4 Playwright workers submitting solutions in the browser.

VERIFIED RESULT

230+

accepted

02

All 15 Repositories

Explore all public codebases across autonomous agents, deep learning, computer vision and analytics.

Showing 15 of 15 repositories

  • RAI-Axiom

    Balamurugan P G and Jason Pandian

    Ongoingfeatured

    A portfolio-allocation policy trained only on synthetic markets, then evaluated zero-shot on six real universes with 10-seed confidence intervals. PPO, no real prices during training. Research by Balamurugan P G and Jason Pandian. Ongoing project; does not claim to beat buy-and-hold (tested tie with SPY out of sample and behind SPY on the holdout point estimate).

    PPOReinforcement LearningSynthetic MarketsSim-to-RealZero-Shot Evaluation
    GitHub
  • VeraScan

    featuredpypi

    Python package that detects train/eval data contamination with 4 checks: exact match, 13-gram overlap, fuzzy match and semantic similarity. Published on PyPI: pip install verascan.

    PythonPyPIExact match13-gram overlapFuzzy match
    GitHub PyPI
  • Bull/Bear Autopilot

    featured

    Multi-agent system for Indian IPO analysis. Every fact is cited to a page, then checked by a deterministic fact-checker; results are delivered on Telegram. 568 tests.

    Multi-agent LLMPage-cited factsDeterministic fact-checkTelegram bot
    GitHub
  • LungScan AI

    featuredlive

    5-class chest X-ray classifier built on DenseNet121, with Grad-CAM heatmaps. 86–88% test accuracy. Full-stack app (FastAPI backend, React frontend, Supabase) deployed on GCP, live at medrag.in.

    DenseNet121Grad-CAMFastAPIReactSupabase
  • PneumoScan

    featuredlive

    Pneumonia detection from chest X-rays with MobileNetV2. 92.5% test accuracy. Live Streamlit app.

    MobileNetV2CNNStreamlit
  • AI LeetCode Agent

    Agent that solves LeetCode problems using a 5-provider LLM fallback chain and 4 Playwright browser workers. 230+ accepted.

    LLM agents5-provider fallbackPlaywright
    GitHub
  • ChurnGuard Pro

    Customer churn prediction dashboard built in Streamlit, trained with Gradient Boosting on 7,043 customers.

    Gradient BoostingStreamlit
    GitHub
  • OncoInsight Pro

    Cancer analytics dashboard in Streamlit over 17,686 records, with statistical analysis plus Gradient Boosting regression and classification (GBR/GBC).

    StreamlitStatisticsGBRGBC
    GitHub
  • Household Power RNN

    Recurrent neural network (RNN) on household power consumption data.

    RNN
    GitHub
  • Titanic ANN

    Artificial neural network (ANN) that predicts Titanic passenger survival.

    ANN
    GitHub
  • Hierarchical Clustering

    Hierarchical clustering on a height–weight dataset.

    Hierarchical clustering
    GitHub
  • DBSCAN Clustering

    DBSCAN density-based clustering on a height–weight dataset.

    DBSCAN
    GitHub
  • K-Means Clustering

    K-Means clustering on a height–weight dataset.

    K-Means
    GitHub
  • Video Game Sales EDA

    Exploratory data analysis of video game sales.

    EDA
    GitHub
  • Luxury Car Dealership

    Luxury car dealership web project.

    Web
    GitHub
03

Experience & Education

Formal computer science education and verified AI/ML engineering internships.

INTERNSHIPS

2 ROLES

Neo Zeno Talent

May–Aug 2026

Intern · Remote

Yuva Intern

Mar–Apr 2026

Intern · Automotive AI · Remote

ACADEMIC BACKGROUND

CGPA 90.6%

B.E. Computer Science and Engineering

2023–2027

Nehru Institute of Technology, Coimbatore (Anna University)

CGPA 90.6% · First in CSE, 2025–2026

Higher Secondary (12th)

Sourashtra HSS, Ramanathapuram

80%

04

Technical Matrix

Applied competencies across the full AI lifecycle from architecture to production evaluation.

🧠

ML & Deep Learning

7 competencies
🤖

LLMs & Agents

5 competencies
🔬

Evaluation & Data

4 competencies
⚙️

Engineering

11 competencies
IMPLEMENTATION RECEIPT
Benchmark Integrity

Train/eval contamination detection

Production audit tool that catches subtle data leakage between training corpora and evaluation benchmarks.

VERIFIED METRIC / PROOFShipped on PyPI (pip install verascan)
SHIPPED IN 1 REPOSITORIES
VeraScan

Python package that detects train/eval data contamination with 4 checks: exact match, 13-gram overlap, fuzzy match and semantic similarity. Published on PyPI: pip install verascan.

Want proof?
05

Certifications (12)

Verified industry and professional credentials in data science, generative AI, and machine learning.

OCI Data Science Professional (Oracle)✓ verified
OCI Generative AI Professional (Oracle)✓ verified
OCI AI Foundations (Oracle)✓ verified
Oracle AI Vector Search (Oracle)✓ verified
IBM Data Science (IBM)✓ verified
Deep Learning (Andrew Ng)✓ verified
Machine Learning with Python✓ verified
Google AI Essentials (Google)✓ verified
Generative AI Leader✓ verified
Linux Fundamentals✓ verified
Tableau✓ verified
NPTEL IoT (NPTEL)✓ verified
06

Grounded Architecture

How the portfolio assistant guarantees factual answers with zero open-web hallucination.

01Deterministic Corpus

Retrieve

Keyword and sectional token matching across all 14 project records, experience, education and resume text. Out-of-domain queries are rejected instantly.

02Bidirectional Grounding

Cite

Every matched document produces interactive citation chips. Users can click any citation chip to scroll directly to the corresponding card on this page.

03Zero Hallucination

Generate & Guard

Single LLM call bounded to context. An anti-hallucination guard rejects any output introducing numbers or academic claims not found verbatim in source docs.

Test the grounded pipeline right now: