I'm Hitesh, and I build scalable AI systems.
Mostly I like the unglamorous half of the work: the part where a promising notebook has to survive real traffic. Training a model is the easy bit — getting inference under 100ms, searching millions of embeddings without falling over, and keeping the whole thing running at 3am is where the actual engineering lives.
- Custom ML models trained and deployed to production, in PyTorch on CUDA
- Inference optimised to sub-100ms at scale with ONNX and TensorRT
- Vector search over millions of embeddings with Milvus and Qdrant
- RAG pipelines, event-driven infra with Kafka, multi-cloud on AWS and GCP with Terraform
I work in Go, Python, and TypeScript, and I care about clear architecture and shipping things people actually use. Always up for talking through a system someone is stuck on.
Languages
AI / ML
CUDA · ONNX · TensorRT · RAG · LLMs · MLOps
Cloud & Infra
Data
Milvus · Qdrant · Pinecone · OpenSearch · Couchbase
Frontend & Mobile
sisara.in · sisarahitesh@gmail.com · Ahmedabad, India




