Job Description
Core Engineering:
- Write and ship production-grade Python (and SQL) that runs reliably against sensitive, high-volume healthcare data.
- Turn models, features, and LLM pipelines from a rough notebook or a senior engineer’s design into scheduled, instrumented production services – and own them once they are live.
- Build and maintain data, model, and application pipelines with the logging, monitoring, and reproducibility that keep them trustworthy in production.
- Build the evaluation that lets the team trust a change before it ships – the eval set, the metrics, the honest before-and-after – offline, and online where the stakes justify it.
- Reproduce, triage, and fix issues in live systems, then add the tests or monitoring that stop them recurring.
- Keep an eye on the cost, latency, and reliability of the models and services you run — production ML and AI carry a real operational bill.
Depending on the focus area, the responsibility will also entail:
- Machine Learning & Modeling — build, evaluate, and productionize models end to end: feature engineering, training, and rigorous offline and online evaluation with drift and fairness monitoring.
- AI / GenAI Engineering — build RAG and agentic pipelines, with the prompt and context engineering, evaluation, and observability that keep them reliable where wrong answers have real consequences.
- ML Platform & MLOps — build feature stores, model deployment and serving, CI/CD for ML, and pipeline orchestration; own the tooling other engineers build on.
- This role may require flexibility across shifts and locations based on business needs.
QUALIFICATION
Minimum Education: B.E./B.Tech in Computer Science, Information Technology or related disciplines.
Candidates should be from IITs, NITs, or other premier engineering institutes only.
All semesters should be cleared with an aggregate CGPA of 8.0 and above.
SKILLS AND COMPETENCIES
Foundation:
- Genuine proficiency in Python, with solid command of data structures, algorithms, and problem-solving – and code you would be willing to show us: organized into functions and modules, version-controlled, with tests, not only notebook cells.
- Comfortable across the full development loop: reading an unfamiliar codebase, writing tests, and shipping changes through review and CI.
- Fluency with SQL and working with real, messy data at some scale.
Demonstrated through projects, internships, research, competitions (Kaggle), or open source:
- At least one ML or AI system you have built that ran outside a notebook — a deployed model, a data or feature pipeline, a working RAG or agent prototype — that you can walk through, including the trade-offs you made and what you would change.
- An understanding that a good offline metric is not a working system: awareness of data leakage, distribution shift, and how you would measure success in production.
- Depth in at least one focus area: Applied ML and evaluation; MLOps and cloud/deployment tooling; or LLMs, RAG, and agentic frameworks.
Bonus — any of this help you stand out:
- Exposure to a major cloud (AWS / GCP / Azure) and containerized deployment.
- Hands-on with orchestration, feature stores, or experiment tracking (Airflow / Kubeflow, Feast, MLflow).
- Practical RAG or agentic work, and familiarity with LLM evaluation, guardrails, and tracing / observability.
- Any exposure to healthcare, insurance, or another regulated, data-sensitive domain.





