Staff Research Engineer/Scientist
ServiceNow · Santa Clara, California, United States
- Senior
- Full-time
- $202,500 – $354,400
- Posted 2026-09-02
- Confirmed live on 25 September 2026
Job description
Job Description
Our Core AI Research team develops novel methods for enterprise agents that reason over multimodal information, use tools, take reliable action across stateful workflows, and improve through feedback. We work across LLM model post-training, agent harnesses, training environments, evaluations, ML, search and reasoning systems, partnering closely with product, engineering, infrastructure, security, and domain experts.
About the role
As a Staff Research Scientist, you will independently lead a major workstream in agent learning and recursive self-improvement. You will turn systematic failures and successful trajectories into hypotheses, experiments, training signals, and deployable improvements to model weights and/or the executable harness around the model.
This is a research role for someone who can move between scientific reasoning, training code, agent systems, and production constraints.
What you get to do in this role:
• Design and execute end-to-end research projects that improve long-horizon enterprise agents across planning, reasoning, memory, tool use, retrieval, computer use, multi-agent coordination, and verification.
• Research model post-training methods such as continued pretraining, supervised fine-tuning (SFT), RL, DPO/GRPO, reward modeling, and distillation.
• Research harness-level optimization across prompts and task framing, tool and schema design, skills, MCP-backed providers, subagents, context and memory management, agent-loop policy, and reliable verifiers.
• Build improvement flywheels that mine trajectories and production-safe signals, identify recurring failure modes, generate or curate data, propose interventions, and measure generalization before promotion.
• Create realistic, stateful training environments and benchmarks for enterprise workflows, with programmatic verifiers and calibrated human or model-based graders where deterministic grading is not possible.
• Run rigorous ablations and scaling experiments; reason explicitly about variance, contamination, reward hacking, distribution shift, cross-model transfer, cost, and latency.
• Develop capabilities across one or more modalities - language, documents, images/video, and speech/audio - and across multilingual or cross-lingual settings.
• Build reproducible distributed pipelines for training, rollout generation, evaluation, and inference; profile and resolve bottlenecks that only appear at scale.
• Partner with other researchers, engineering, and product teams to move validated methods into reliable enterprise systems.
• Communicate results through research reviews, technical reports, publications, patents, open-source contributions, and decision-ready recommendations
Qualifications
To be successful in this role you have:
• Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
• 7+ years of relevant ML/AI research or engineering experience, or equivalent evidence of research depth and impact. A PhD or other advanced degree is required.
• A track record of independently taking ambiguous research questions from hypothesis through implementation, experiment, analysis, and measurable outcome.
• Strong foundations in machine learning, deep learning, reinforcement learning and/or probabilistic/statistical experimentation, with hands-on experience training or adapting large language models or multi-modal models.
• Advanced Python and PyTorch skills, including the ability to modify training loops, model code, data pipelines, evaluators, or research infrastructure.
• Practical depth in agentic AI: tool use, environments, planning/reasoning, memory/context, retrieval, long-horizon execution, or multi-agent systems.
• Experience designing decision-useful evaluations, including strong dataset partitions, trajectory analysis, graders/verifiers, error taxonomies, and robust conclusions under noisy measurements.
• Experience with distributed training, rollout, and/or inference using some subset of PyTorch distributed/FSDP, DeepSpeed or Megatron; verl, TRL, OpenRLHF or comparable post-training systems; and vLLM, SGLang, or comparable inference stacks.
• Strong software engineering fundamentals and the ability to build reliable, testable, reproducible systems in collaboration with AI/ML infra and software engineers.
• Publications at top-tier venues (ICLR, NeurIPS, ICML, ACL, EMNLP, AAAI).
Preferred qualifications:
• Research and production experience in multimodal, document AI, computer vision, speech/audio, multilingual, or cross-lingual modeling.
• Experience with enterprise agents, stateful workflows, browser/computer use, MCP or similar tool protocols, simulation environments.
• Experience with synthetic data, model-generated feedback, prop
Interview problems reported for ServiceNow
Reported by candidates and public write-ups, not by ServiceNow. Practise each one here:
- Longest Substring Without Repeating Characters — Medium
- Number of Islands — Medium
- Container With Most Water — Medium
- Longest Repeating Character Replacement — Medium
- Valid Parentheses — Easy
- Two Sum — Easy
- Merge Two Sorted Lists — Easy
- Longest Palindromic Substring — Medium
- Coin Change — Medium
- Maximum Subarray — Medium
- Set Matrix Zeroes — Medium
- Group Anagrams — Medium
- Best Time to Buy and Sell Stock — Easy
- Top K Frequent Elements — Medium
- Product of Array Except Self — Medium
- Reverse Linked List — Easy
- Combination Sum — Medium
- Pacific Atlantic Water Flow — Medium
- House Robber II — Medium
- Longest Common Subsequence — Medium
- Merge Intervals — Medium
- Maximum Product Subarray — Medium
More at ServiceNow
- Senior Analyst, US International Tax · Santa Clara, CALIFORNIA, United States
- Advisory Solution Consultant - State Government · Sydney, NSW, Australia
- Senior in-Market Engineer · Tokyo, , Japan
- Principal Software Engineer · San Diego, CALIFORNIA, United States
- Staff Software Engineer · Santa Clara, CALIFORNIA, United States
- Staff Software Engineer · San Diego, CALIFORNIA, United States
- Principal Software Engineer · Santa Clara, CALIFORNIA, United States
- Pricing Operations Senior Analyst · Remote; San Francisco de Heredia, Heredia, Costa Rica