AI / Machine Learning Software Engineer

I build reliable AI systems from retrieval to GPU kernels.

I’m Ranjith Vutnoor, an IIT Bhilai alumnus working across production RAG, LLM evaluation, agentic workflows, model fine-tuning and GPU-accelerated machine learning.

Open to thoughtful AI/ML opportunities Hyderabad, India 2.5+ years in production AI systems
240+engineering work items delivered across features, reliability and defects
12.68×encoder speed-up in CUDA-accelerated thesis experiments
0.53→0.668Nemotron reasoning challenge score improvement
240+algorithm problems solved while strengthening interview fundamentals

Selected work / 01

Systems designed around evidence, evaluation and engineering constraints.

The public case studies below are intentionally sanitised. They show the technical approach without exposing confidential data, internal code or client-specific architecture.

Production AIRAG · Retrieval

Enterprise RAG and semantic search

A grounded question-answering system spanning query understanding, expansion, dense and hybrid retrieval, reranking, evidence selection, citations and reflective retries.

Quality systemsEvaluation

EvalStudio — LLM quality workspace

An evaluation workflow for repeatability, golden-set regression, LLM-as-judge scoring and release-oriented quality gates.

IIT Bhilai thesisCUDA · Industrial AI

GPU-accelerated anomaly detection

Extended MSCRED and ConvLSTM for industrial time-series anomalies, then accelerated convolution through im2col, GEMM, shared memory and a custom PyTorch CUDA extension.

NVIDIA · KaggleFine-tuning

Nemotron structured reasoning

Built compact synthetic reasoning traces, used a 4-bit QLoRA-style workflow, resolved model compatibility issues and improved benchmark accuracy.

Experience / 02

Building depth across the AI application stack.

I work best where research ideas must survive real data, production constraints and measurable quality expectations.

2023 — present

AI / Machine Learning Software Engineer

HCLTech · Enterprise AI systems

Developing and maintaining a production knowledge assistant for complex technical documentation.

  • Engineered query classification, expansion, vector and hybrid retrieval, reranking and grounded response generation.
  • Improved reliability through citation validation, guardrails, retry logic, observability and regression testing.
  • Supported model migrations across multiple OpenAI model generations and structured response workflows.
  • Designed incremental metadata-driven ingestion to avoid unnecessary full-index rebuilds.
2021 — 2023

M.Tech — Data Science and Artificial Intelligence

Indian Institute of Technology Bhilai

Built a research foundation across machine learning, deep learning, data science, optimisation and GPU computing.

  • Completed a thesis on industrial multivariate time-series anomaly detection.
  • Worked with MSCRED, ConvLSTM, PyTorch, CUDA, im2col and GEMM optimisation.
  • Benchmarked CPU, global-memory GPU and shared-memory GPU implementations.
Independent work

Applied research and product experiments

Kaggle · Local LLMs · Desktop software

Exploring model adaptation, local inference, reasoning datasets and human-centred AI product ideas.

  • Fine-tuned a LoRA adapter for the NVIDIA Nemotron reasoning challenge.
  • Compared quantised language and vision-language models through local inference tools.
  • Prototyped an Electron desktop companion with productivity and Pomodoro workflows.

Expertise / 03

A practical stack for modern AI engineering.

I care about the complete path from source data to user-facing reliability—not only the model call.

01 · AI systems

RAG and agents

Query understanding, expansion, embeddings, hybrid search, reranking, reflective retrieval, citations, structured outputs and tool integration.

02 · Evaluation

Quality engineering

MRR, NDCG, hit rate, relevance labels, golden datasets, LLM-as-judge, repeatability, regression gates and failure analysis.

03 · Model adaptation

Fine-tuning and inference

LoRA, QLoRA-style workflows, PEFT, synthetic reasoning data, quantisation, checkpoint selection and local model evaluation.

04 · Performance

GPU-accelerated ML

CUDA kernel programming, im2col, GEMM, shared-memory tiling, custom PyTorch extensions, ConvLSTM and performance benchmarking.

05 · Engineering

Application development

Python, SQL, PostgreSQL, APIs, React, Electron, Git, testing, data processing and production troubleshooting.

06 · Platforms

Cloud and observability

Azure OpenAI, Azure App Service, Key Vault, Application Insights, retries, diagnostics and operational feedback loops.

07 · ML frameworks

Research tooling

PyTorch, Hugging Face, PEFT, TRL, bitsandbytes, Qdrant, Jupyter, Ollama and LM Studio.

08 · Foundations

Algorithms and systems

Data structures, algorithmic problem-solving, system design, distributed-system concepts and model-serving architecture.

Beyond work / 05

Curiosity beyond a single technology stack.

Product and technology exploration

I enjoy evaluating local AI models, personal computing, developer tools and product concepts—especially where engineering decisions must balance quality, latency, privacy and usability.

Travel and cultural exploration

Travel, temples and heritage destinations offer a useful counterweight to technical work. They encourage patience, observation and a longer view of systems and human behaviour.

Contact / 06

Let’s build something rigorous and useful.

I’m interested in AI/ML engineering roles and technical conversations involving production RAG, evaluation, LLM systems, model adaptation and GPU-accelerated machine learning.