Senior Machine Learning Engineer
Software Engineering
United States
USD 130k-160k / year
Position Summary
Candid is a nonprofit that provides the most comprehensive data and insights about the social sector. We get you the information you need to do good. Candid currently has an opportunity for a Senior Machine Learning Engineer. Candid’s Data Science and AI team builds and ships machine learning and AI models across many concurrent projects. The Senior ML Engineer exists to take operational ownership of Candid’s deployed ML and AI services to free up data science capacity. This is primarily an operational ownership role, with a secondary focus on incrementally improving deployment practices, observability, and inference performance as that ownership surfaces the need. The right person is motivated by reliability, ownership, and enabling others to do their best work; brings cost-awareness as a default in a budget-conscious nonprofit; and is energized by building standardized practices. The Senior ML Engineer will work with data scientists and data engineers to deploy services at scale, and collaborate on cross-functional teams to understand product needs.
Position: Senior Machine Learning Engineer
Reporting to: Data Science Manager
Supervises: N/A
Schedule: 35-hour work week, Monday through Friday
Compensation: $130,000 - $160,000 (this range is for the NYC area and will be adjusted for other localities; additionally, factors like skills and experience will be considered).
Location: Remote. In-person attendance is expected twice per year during our annual, weeklong all-staff summits. Additional in-person meeting participation is expected at least once per quarter for senior leaders and at least once per month for the executive team. Staff not located in the NYC area are expected to travel for these meetings.
Benefits: Health insurance (medical, dental, vision), retirement contribution with additional option for a match, paid life insurance and AD&D, paid leave time (PTO, compassionate leave, volunteer, holiday, parental), short-term and long-term disability, pre-tax transit, flexible spending accounts, supplemental insurance, summer hours, and Public Service Loan Forgiveness (PSLF) program eligible employer.
Responsibilities
- Take operational ownership of Candid’s deployed machine learning and AI services — monitoring for degradation, managing retraining cadences, coordinating handoffs from data scientists, and serving as the accountable point of contact for models in production.
- Improve inference performance for deployed models, including complex graph inference models, applying techniques such as quantization, artifact slimming, batching, and efficient serialization.
- Design and operate experiment tracking, model versioning, and artifact management so data scientists have a consistent, low-friction way to hand off work.
- Build and maintain observability for ML and AI services — centralized logging, metrics, dashboards, and alerting for the services that matter most — and proactively reduce incidents through good deployment hygiene.
- Establish and iteratively improve a repeatable deployment path for new ML services on AWS, including CI/CD integration and infrastructure-as-code patterns.
- Monitor and actively manage AWS spend across ML workloads, including Bedrock token usage, compute sizing, S3 lifecycle, and publish cost visibility to the team.
- Work closely with data scientists to understand model behavior, surface operational insights back to them, and translate research code into production-ready deployments.
- Serve as the primary technical liaison between the Data Science team and Candid’s product and software engineering teams when ML or AI services are integrated into products and systems, defining integration contracts, APIs, latency and reliability expectations, and input/output schemas, and supporting those teams through integration.
- Build and operate Amazon Bedrock-backed services and integrations where they are the right tool for the job.
- Contribute to Candid’s secure-by-default posture for ML services, including IAM scoping, secrets handling, and compliance-aligned tagging.
- Participate in technical planning and roadmap discussions for the Data Science team.
Requirements
- 4+ years of professional software engineering experience, including at least 2 years in a role where production ML systems were a primary responsibility (MLOps, ML engineering, or production-ML-focused data science).
- Strong proficiency in Python, including writing production-quality service code.
- Hands-on, operational experience with experiment tracking and model lifecycle tooling in production, such as MLflow, Weights & Biases, or comparable systems.
- Hands-on experience deploying PyTorch models in production, including familiarity with serving optimization techniques such as quantization, batching, ONNX Runtime, model serialization, container/artifact slimming, or cold-start mitigation on Lambda/Fargate.
- Demonstrated experience deploying and monitoring ML models in production, with an understanding of model degradation, drift signals, retraining triggers, and artifact management.
- Working knowledge of deploying Python-based ML services on AWS including Lambda, ECS/Fargate, S3, IAM, and CloudWatch. Deep infrastructure expertise is not required; comfort with the AWS ML deployment stack is.
- Experience building or operating CI/CD pipelines for ML services.
- A demonstrated track record of improving production reliability through observability and disciplined deployment and maintaining healthy systems over time.
- Ability to work closely with data scientists, understanding their outputs, inheriting their work, and communicating operational decisions back to them clearly.
- Demonstrated ability to work cross-functionally with software or product engineering teams, translating ML service capabilities into integration-ready contracts and supporting those teams through adoption.
- Comfort owning work independently.
- Strong written and verbal communication.
- Willingness to perform other duties and special projects as needed/requested.
- Sensitivity and respect for racial, gender, sexual orientation, and cultural differences.
- Commitment to Candid's values: driven, direct, accessible, curious, and inclusive.
Preferred qualifications
- Production graph ML inference experience (PyTorch Geometric, DGL, or comparable).
- Experience with optimized CPU inference runtimes (ONNX Runtime, TorchScript) or serving frameworks.
- Experience with Amazon Bedrock or another managed LLM service (Anthropic API, Azure OpenAI, Vertex), and/or operating RAG systems or agentic AI workflows.
- Familiarity with infrastructure as code (CloudFormation, CDK, Terraform, SAM).
- Experience with high performance programming languages for ML serving (Rust, Go, etc.) is a nice to have.
- Experience at a nonprofit or mission-driven organization, and mentoring data scientists on production engineering practices.
About Candid
Candid’s mission is to get you the information you need to do good.
The world’s problems are only growing, and change can’t wait. Nonprofits are needed now more than ever, but all too often their work goes without adequate support.
Candid makes it easier and faster for nonprofits and funders to connect in pursuit of solutions to change the world. Candid is where nonprofits find grants, donors find nonprofits that inspire them, and all can gain insights about the work being done for good.
Candid is a qualifying nonprofit organization as defined by the Public Service Loan Forgiveness Program. As such, Candid employees may claim their employment time on their PSLF application. We offer a competitive salary and excellent benefits. Due to the high volume of applicants we typically receive, we regret that we can only contact candidates we would like to interview.
For more information on positions available at Candid, please visit our website: Work with us
Candid is an equal opportunity employer. Candid provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local laws.
This policy applies to all terms and conditions of employment, including recruiting, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation, and training.