Enterprise RAG and semantic search
A grounded question-answering system spanning query understanding, expansion, dense and hybrid retrieval, reranking, evidence selection, citations and reflective retries.
AI / Machine Learning Software Engineer
I’m Ranjith Vutnoor, an IIT Bhilai alumnus working across production RAG, LLM evaluation, agentic workflows, model fine-tuning and GPU-accelerated machine learning.
Selected work / 01
The public case studies below are intentionally sanitised. They show the technical approach without exposing confidential data, internal code or client-specific architecture.
A grounded question-answering system spanning query understanding, expansion, dense and hybrid retrieval, reranking, evidence selection, citations and reflective retries.
An evaluation workflow for repeatability, golden-set regression, LLM-as-judge scoring and release-oriented quality gates.
Extended MSCRED and ConvLSTM for industrial time-series anomalies, then accelerated convolution through im2col, GEMM, shared memory and a custom PyTorch CUDA extension.
Built compact synthetic reasoning traces, used a 4-bit QLoRA-style workflow, resolved model compatibility issues and improved benchmark accuracy.
Experience / 02
I work best where research ideas must survive real data, production constraints and measurable quality expectations.
HCLTech · Enterprise AI systems
Developing and maintaining a production knowledge assistant for complex technical documentation.
Indian Institute of Technology Bhilai
Built a research foundation across machine learning, deep learning, data science, optimisation and GPU computing.
Kaggle · Local LLMs · Desktop software
Exploring model adaptation, local inference, reasoning datasets and human-centred AI product ideas.
Expertise / 03
I care about the complete path from source data to user-facing reliability—not only the model call.
Query understanding, expansion, embeddings, hybrid search, reranking, reflective retrieval, citations, structured outputs and tool integration.
MRR, NDCG, hit rate, relevance labels, golden datasets, LLM-as-judge, repeatability, regression gates and failure analysis.
LoRA, QLoRA-style workflows, PEFT, synthetic reasoning data, quantisation, checkpoint selection and local model evaluation.
CUDA kernel programming, im2col, GEMM, shared-memory tiling, custom PyTorch extensions, ConvLSTM and performance benchmarking.
Python, SQL, PostgreSQL, APIs, React, Electron, Git, testing, data processing and production troubleshooting.
Azure OpenAI, Azure App Service, Key Vault, Application Insights, retries, diagnostics and operational feedback loops.
PyTorch, Hugging Face, PEFT, TRL, bitsandbytes, Qdrant, Jupyter, Ollama and LM Studio.
Data structures, algorithmic problem-solving, system design, distributed-system concepts and model-serving architecture.
Writing / 04
Long-form articles are being prepared for this site and DEV Community. The outlines are available now.
From convolution bottleneck to a custom PyTorch extension and shared-memory performance analysis.
In preparationCompact synthetic traces, 4-bit adaptation, compatibility fixes and deterministic evaluation.
In preparationQuery expansion, hybrid retrieval, reranking, reflective search and measurable relevance.
Beyond work / 05
I enjoy evaluating local AI models, personal computing, developer tools and product concepts—especially where engineering decisions must balance quality, latency, privacy and usability.
Travel, temples and heritage destinations offer a useful counterweight to technical work. They encourage patience, observation and a longer view of systems and human behaviour.
Contact / 06
I’m interested in AI/ML engineering roles and technical conversations involving production RAG, evaluation, LLM systems, model adaptation and GPU-accelerated machine learning.