- Experiencia
- 3+ años
- Salario
- USD 140,000 – USD 165,000 / year
- Vacantes
- 1
- Al corriente
- hace 1 hora
- Modo de trabajo
- En la oficina
- Educación
- Bachelor's or Master's in Data Science, Computer Science, Statistics, Mathematics, or related field
- Reanudar
- Se requiere solicitud
Descripción del trabajo
About the Role
ScienceLogic is transforming IT operations by enabling autonomic capabilities where systems self-heal, self-optimize, and align seamlessly with business goals. Their platform offers unified insights across hybrid and multi-cloud environments, automating workflows and driving scalable performance. They seek a talented Data Scientist to strengthen their Data Science team working with a unique architecture of multiple small, local language models (LLMs) instead of relying on large, hosted frontier APIs. This role focuses on maximizing outcomes through evaluation, routing, prompting, grounding, and orchestration of these models.
Key Responsibilities
- Design comprehensive evaluation frameworks for LLM and agent outputs, including golden datasets, regression tests, and rubric-based scoring systems.
- Develop and fine-tune LLM-as-judge pipelines, validating them against human-labeled data while managing their biases and variances.
- Establish and monitor response-quality metrics such as groundedness, hallucination frequency, answer relevance, completeness, instruction adherence, and persona consistency.
- Maintain and expand evaluation datasets as the product evolves, benchmarking models to allocate tasks and quantify trade-offs between smaller local models and larger alternatives.
- Conduct adversarial testing including prompt injections, jailbreak scenarios, misuse detection, and edge-case identification to ensure model robustness.
- Implement chaos and stress tests to evaluate model reliability under degraded or hostile environments, identifying failure modes and feeding insights into safety guardrails.
- Assess retrieval effectiveness on document corpora with precision metrics and experiment with indexing and hybrid retrieval strategies.
- Analyze multi-step agent behaviors including tool call accuracy, trajectory efficiency, state replayability, and guardrail violations.
- Monitor and evaluate intent classification and routing accuracy as measurable phenomena rather than opaque systems.
- Build continuous evaluation systems to detect regressions upon model changes, prompt updates, or pipeline modifications while discriminating genuine quality drifts from stochastic noise.
- Recommend and verify corrective actions such as prompt adjustments, retrieval enhancements, routing changes, or model switches.
- Create and deploy production-grade predictive models to forecast operational trends, detect anomalies, signal early warnings, and support capacity planning based on telemetry data.
- Oversee model deployment lifecycle including monitoring, recalibration, and retraining aligned with shifting data and behaviors.
- Define operationally relevant metrics like precision/recall on incident predictions, forecast errors, and lead time for early signals.
- Integrate predictive outputs into LLM reasoning layers to influence advisories and operator recommendations.
- Apply domain analytics in AIOps/NOC such as log anomaly detection, event correlation, root-cause analysis, and performance metrics linking token usage and costs to customer value indicators like MTTR and operator labor.
- Communicate analytical results clearly to engineering and product teams providing actionable insights.
- Leverage LLM-assisted workflows to scale evaluation efforts including analysis drafting, synthetic case generation, and labeled data bootstrapping.
- Continuously adopt cutting-edge evaluation, retrieval, and agentic-analysis techniques enhancing team methodologies.
Requirements
- Bachelor's or Master's degree in Data Science, Computer Science, Mathematics, Statistics, or a related discipline; or equivalent professional experience.
- Minimum three years of practical experience in data science, machine learning, or quantitative analytics.
- Expertise in applied statistics with strong judgment to design meaningful experiments and significance tests for noisy, non-deterministic system outputs.
- Proven track record building, deploying, and maintaining production predictive or time-series models focusing on forecasting, anomaly detection, trend analysis, and recalibration.
- Hands-on experience evaluating, analyzing, or enhancing large language models or NLP systems through evaluation design, quality measurement, or retrieval assessment.
- Proficient in Python programming and advanced SQL skills capable of queries on large analytical datasets.
- Comfortable with foundational model concepts and practical engagement with modern LLM tooling layers including eval/harness frameworks, judge pipelines, serving, prompting, and testing libraries.
- Adept at building analysis and visualization solutions programmatically.
Preferred Qualifications
- Experience optimizing the performance of small or locally hosted models within compute, memory, or latency constraints using quantization-aware evaluation, prompt/context tuning, or model routing strategies.
- Knowledge of retrieval-augmented systems and large-scale retrieval evaluation methodologies.
- Familiarity with agentic frameworks involving tool orchestration, human-in-the-loop processes, and replayable-state management.
- Background in adversarial robustness and red-teaming techniques focused on LLM systems.
- Domain expertise in IT operations management including AIOps, Network Operations Centers (NOC), IT Service Management (ITSM), observability, and telemetry-based anomaly detection.
- Experience working with large-scale analytic and big data platforms.
- Exposure to cloud technologies supporting data science and machine learning workflows.
- Understanding of enterprise security and compliance requirements in product delivery contexts.
Benefits and Perks
- Comprehensive healthcare coverage encompassing medical, dental, and vision plans.
- 401(k) retirement plan featuring employer matching contributions.
- Flexible Paid Time Off policy to support personal rejuvenation and work-life balance.
- Volunteer Time Off granting two days annually to engage with charitable organizations of choice.
- Service milestone sabbatical every five years of employment.
- Paid parental leave to support family growth.
- Attractive employee referral bonus program.
- Pet insurance options available.
- Access to a central office in Reston Town Center with amenities including stocked kitchen, rotating snacks and beverages, and weekly catered lunches.
- Regular virtual community events such as cooking classes, yoga, and meditation sessions.
- Opportunities for professional development with industry-leading experts.
Equal Opportunity Statement
ScienceLogic is committed to creating a diverse and inclusive workplace. Applicants are encouraged to apply even if they do not meet every qualification, recognizing that diverse candidates contribute significantly. Employment decisions are made regardless of race, color, religion, gender, sexual orientation, gender identity, national origin, or legally protected characteristics relevant to the application location.
Compensation
The salary range for this position is $140,000 to $165,000 annually.