CFO-Copilot
Autonomous financial agent with SmolAgent, RAG over ERP/PostgreSQL, Qdrant, Redis caching, and Chronos-2/N-Hits/Prophet forecasting.
Hey There, This is
AI & ML Engineer · Bengaluru, Karnataka
And I'm an
AI & ML Engineer based in Bengaluru, Karnataka, currently building autonomous agent systems at Thrivv AI (Remote, Dubai). I specialize in machine learning engineering (feature pipelines, model training, MLflow, drift monitoring), MCP-powered agents, hybrid RAG, LLM fine-tuning, eval-driven LLMOps, and production inference pipelines — from multimodal interview systems to financial copilots and document forensics engines.
Previously at Vacanzi as AI Engineer and at Rubixe AI Solution as a Data Scientist Consultant, I've delivered classical ML models, forecasting systems, RAG chatbots, computer vision pipelines, and client analytics solutions. I work across PyTorch, Scikit-Learn, XGBoost, LightGBM, MLflow, Airflow, LangChain, LangGraph, SmolAgent, MCP, vLLM, FastAPI, Qdrant, Redis, PostgreSQL, and AWS — with a focus on evaluated, guarded, low-latency production AI & ML systems.
Feature pipelines, XGBoost/LightGBM, MLflow registry, drift monitoring, and batch/online inference.
MCP tool agents, hybrid RAG, and LoRA-tuned models built for real workloads.
Evals, guardrails, and observable ML/LLM pipelines aligned to business outcomes.
I combine deep technical expertise in machine learning engineering, agentic AI, hybrid RAG, and production LLMOps with clear communication — translating complex engineering into business outcomes. I invest time upfront to understand the problem, then build evaluated, guarded, scalable systems that perform in production.
Autonomous financial agent with SmolAgent, RAG over ERP/PostgreSQL, Qdrant, Redis caching, and Chronos-2/N-Hits/Prophet forecasting.
Fine-tuned Mistral 7B v0.3 with PEFT (LoRA, QLoRA), custom preprocessing pipelines, and domain-specific evaluation.
Coqui XTTS v2 fine-tuned on Indian accent data with Librosa/Whisper preprocessing and ONNX inference optimization.
Digital forensics engine using DCT, FFT, ELA, EXIF validation, PDF analysis, and Llama 3 explainability reports.
Real-time multimodal video interviews with LangChain agents, OpenCV/TensorFlow CV, and audio communication analysis.
Voice-enabled RAG chatbot with TTS/STT pipelines and MongoDB for live recruiter job data.
SQL analytics on IMDB movies and HR dashboards in Power BI / Tableau for decision-ready insights.
Design end-to-end ML systems — feature engineering, XGBoost / LightGBM / Scikit-Learn / PyTorch models, Optuna tuning, offline evaluation, and online or batch inference at production scale.
Build continuous training pipelines with Airflow, Spark, Kafka, Ray, MLflow model registry, feature stores, drift monitoring, A/B tests, and CI/CD for reliable model rollout.
Design and deploy autonomous LLM agents with MCP tools, SmolAgent, LangGraph, structured outputs, multi-step reasoning, and secure local or cloud hosting for enterprise workflows.
Build production RAG and CRAG with hybrid BM25 + dense search, embeddings design, Qdrant or pgvector, Cohere reranking, Redis caching, and live ERP / PostgreSQL / MongoDB sources.
Fine-tune open-source LLMs (Mistral, Llama) with PEFT, LoRA, and QLoRA on Hugging Face — including preference-style alignment workflows and domain adapters.
Ship eval-driven AI with Ragas / DeepEval harnesses, LangSmith tracing, W&B experiment tracking, regression suites for agents, and cost / latency monitoring.
Serve models at scale with vLLM, ONNX, AWQ / GPTQ / bitsandbytes quantization, streaming APIs, FastAPI, Docker, Kubernetes, and AWS GPU deployments.
Deliver CFO copilots and financial agents with ERP integration, banking data RAG, and time-series forecasting using Chronos-2, N-Hits, and Prophet.
Build real-time CV and VLM pipelines with OpenCV, YOLO, MediaPipe, TensorFlow, and vision-language models for video interviews, behavior analysis, and document vision.
Build conversational voice assistants with Whisper STT, Coqui XTTS custom accent fine-tuning, Librosa audio pipelines, and low-latency duplex UX.
Detect document tampering using DCT, FFT, ELA, EXIF/XMP validation, PDF structural analysis, and Llama-powered explainability reports.
Add production guardrails — moderation, PII redaction, prompt-injection defenses, policy filters, and audit-friendly logging for enterprise AI assistants.
Integrate OpenAI, Anthropic, Hugging Face, and AWS Bedrock into products with robust prompting, structured outputs, streaming, guardrails, and cost-optimized architectures.
| Institution | Degree | Year | Grade / Status |
|---|---|---|---|
| Indian Institute of Technology Patna | Master of Computer Application (MCA) | Pursuing | CPI: 9.35/10 · SPI: 9.35/10 |
| Bengaluru City University | Bachelor of Computer Application (BCA) | 2023 | First Class |
Available for freelance and contract AI / ML engineering — ML pipelines, agentic systems, hybrid RAG, LLM fine-tuning, evals, guardrails, and production inference. Based in Bengaluru — open to remote worldwide.
Chat about Shashi's experience, skills, projects, services, or how to collaborate on AI engineering work.