- Experience
- 1+ yrs
- Salary
- —
- Openings
- 1
- Posted
- há 1 hora
- Work mode
- In office
- Education
- Master's degree
- Resume
- Required to apply
Where you'll work
Job description
About the Role
IMX Data is seeking a Senior Data Scientist to join their team in Ajmer, Rajasthan. This role involves working extensively with large-scale healthcare claims data to extract insights and develop scalable analytical models and deployment pipelines.
Key Responsibilities
- Develop diagnostic and procedural indexes on multi-terabyte healthcare claims data to detect patterns and cost contributors using SQL, R, and standardized healthcare coding systems like ICD, CPT, and HCPCS.
- Design and implement data schemas in Snowflake with automation using dbt, ensuring strict role-based access controls and cost management via warehouse monitoring.
- Conduct advanced time series analyses on medical and financial data using ARIMA, SARIMAX, and Facebook Prophet models within Jupyter Notebooks to identify seasonal trends and anomalies.
- Build retrieval augmented generation (RAG) pipelines that semantically extract relevant document segments through ChromaDB and Pinecone by applying tokenization, embedding-based cosine similarity, hybrid, and tree search techniques powered by the OpenAI API.
- Deploy custom large language model solutions into production environments utilizing containerized inference pipelines with Docker, FastAPI, and Kubernetes orchestration.
Required Qualifications and Skills
- Master's degree in Computer Science, Computer Information Systems or a related field.
- Minimum one year of professional experience in data science with a focus on healthcare or large-scale data analytics.
- Proficiency in SQL, R, and working knowledge of healthcare data standards (ICD, CPT, HCPCS).
- Experience with Snowflake data warehouse and dbt for data modeling and pipeline automation.
- Familiarity with time series analytical methods and tools including ARIMA, SARIMAX, and Facebook Prophet.
- Hands-on expertise in semantic search techniques using ChromaDB, Pinecone, tokenization, vector similarity calculations, and OpenAI API integration.
- Skilled in container technologies (Docker), API frameworks (FastAPI), and cloud-native orchestration platforms (Kubernetes).