Technical Director, Large-Scale AI Model Inferencing

Samsung Semiconductor · San Jose, California, United States

  • Senior
  • Full-time
  • $219,000 – $351,000
  • Posted 2026-09-09
  • Confirmed live on 25 September 2026

Apply at Samsung Semiconductor

Job description

Please Note:

To provide the best candidate experience amidst our high application volumes, each candidate is limited to 10 applications across all open jobs within a 6-month period.

Advancing the World’s Technology Together

Our technology solutions power the tools you use every day--including smartphones, electric vehicles, hyperscale data centers, IoT devices, and so much more. Here, you’ll have an opportunity to be part of a global leader whose innovative designs are pushing the boundaries of what’s possible and powering the future.

We believe innovation and growth are driven by an inclusive culture and a diverse workforce. We’re dedicated to empowering people to be their true selves. Together, we’re building a better tomorrow for our employees, customers, partners, and communities.

What You’ll Do

Inference is becoming a memory-bandwidth business. As models scale past what any single GPU can hold — KV caches grow with context, MoE expert weights spill beyond HBM, and new architectures change the rules of what "model state" even means — the winners will be the companies that treat memory as the core product of AI inference, not an afterthought.

We are looking for a Hands-on Principal Engineer who combines deep, first-principles knowledge of AI model architectures (dense Transformers, Mixture-of-Experts, State Space Models, and hybrids) with production-scale inference expertise, to own the requirement for full-stack AI memory solutions at scale — spanning GPU HBM, host DRAM, CXL-attached memory pools, and NVMe/SSD tiers and Samsung Cognos, AI memory software that moves model state intelligently across them.

This person will be the technical authority who connects model behavior to memory-system design: someone who can explain why an MoE router's activation pattern dictates an LRU expert cache policy, why a Mamba state cache breaks the assumptions of PagedAttention, and why disaggregated prefill/decode changes the required memory bandwidth per token by an order of magnitude — and then build the products that exploit those facts.

Level: Principal Engineer
Team: Memory Solutions Lab / Data Fabric Solutions
Reports to: Chief Technologist, Memory Solutions Lab

Location: Daily onsite presence preferred at our San Jose office/headquarters in alignment with our Flexible Work policy; remote/hybrid option available.

Job ID: 43027

Model Architecture Expertise — The Foundation

• Serve as expert on how different model families consume and move memory, and translate that into memory-product requirements:

• Dense Transformers: MHA/MQA/GQA/MLA attention, KV-cache growth characteristics, long-context behaviors, attention sinks and prefix locality.

• Mixture-of-Experts: routed vs. shared experts, expert-parallel execution, routing skew and hot-expert locality, expert-weight offloading and cache-admission policies, per-token weight-read economics.

• State Space Models (Mamba/Mamba-2) and hybrid SSM-attention architectures: recurrent state vs. KV cache semantics, state size per sequence and per layer, cache-swapping behavior for context switching and batching, and what "cache-aware scheduling" means when the state is a fixed-size tensor instead of a token-indexed table.

• Emerging architectures: linear attention, sliding-window/hybrid layers, diffusion and multimodal transformers — and how each changes the memory hierarchy math.

• Model the memory footprint, bandwidth demand, and access patterns of frontier open-weight models (e.g., Llama/Qwen-class dense, DeepSeek/Kimi-class MoE, Jamba-class hybrids) and publish internal reference architectures for each.

• Track the model landscape as a roadmap input: anticipate what coming architectures (longer contexts, agentic multi-session reuse, reasoning-loop workloads, speculative decoding drafts) will demand from memory systems 12–24 months out.

Large-Scale Inference Expertise

• Own deep expertise in production inference stacks — SGLang (HiCache), vLLM (PagedAttention, LMCache integration), NVIDIA Dynamo, TensorRT-LLM, llama.cpp-class engines — including their memory-management internals, not just their flags.

• Drive inference performance engineering: continuous batching, chunked prefill, disaggregated prefill/decode, prefix and radix caching, speculative decoding, CUDA Graphs, and their interactions with memory tiering.

• Own the latency/throughput/cost envelope: TTFT and TBT/TPOT SLOs, tokens-per-second per dollar, GPU memory utilization as the binding constraint, and the tradeoff curves between cache hit rate, memory capacity, and bandwidth.

• Define benchmarking and characterization methodology: realistic agentic and long-context workloads (multi-turn reuse, session persistence, RAG prefixes), KV-cache reuse-rate measurement, and bandwidth-latency profiling across the full hierarchy (Nsight, PyTorch Profiler, vendor memory tools).

Full-Stack AI Memory Solutions — The Core Mandate

• Define engineering requirements, with proof, for tiered memory syste

Prepare for the interview

Nothing collected for this employer yet. The Blind 75 is what technical screens draw from; practise it here, with a coach, in Java or Python.

More at Samsung Semiconductor

All open software jobs