LLM inference and generative AI
CUDA or Triton kernels
Distributed inference
Documented ML project
Time-series portfolio evidence
Five or more years in performance-critical systems
Newly verified LLM-inference systems role covering MoE, SSMs, quantization, CUDA/Triton and distributed deployment.
Sponsorship is not stated on the official posting.
The official NVIDIA requisition explicitly requires five or more years.
The employer did not state compensation. No estimate is presented as official.
LLM inference and generative AI
CUDA or Triton kernels
Distributed inference
Documented ML project
Time-series portfolio evidence
Five or more years in performance-critical systems
Research-to-production scope
Research-to-industry portfolio
None material
Employer compensation not stated
Official source captured
Targeted CV draft available
Israel work authorization
5+ years of performance-critical software engineering
English-working market preference
None material