All jobs

Software Development Engineer II, AWS SageMaker AI

Amazon.com Services LLC3h ago
United StatesOnsiteFull-timeMid Level2+ yrs exp
H-1B verified · 2310 LCAs

Top focus

Software EngineerSoftware Engineer IiSenior Software EngineerAws Engineer
  • At AWS SageMaker AI, we're making it easy to build state-of-the-art foundation models on the cloud. Model Factory is our platform for building, training, customizing
  • evaluating foundation models at scale. Instead of hand-chaining data prep, distributed training, evaluation
  • deployment across thousands of GPU and AWS Trainium devices, Model Factory lets teams express the whole lifecycle as a single, contract-validated workflow — orchestrated, reproducible
  • fully managed. As LLMs and Generative AI scale, Model Factory is the platform that turns frontier training research into a reliable, repeatable pipeline for our internal teams and customers. We're looking for a Software Development Engineer to help design, build
  • operate the distributed systems at the core of this platform — the orchestration engine, compute integrations, SDK and contract layer
  • the infrastructure that runs large-scale training and customization jobs. You'll own components end-to-end, from design through delivery and on-call operations
  • work closely with the ML scientists and platform teams who depend on Model Factory every day. You'll turn requirements into robust, scalable, supportable services that fit cleanly into the overall architecture, uphold a high engineering bar
  • grow your scope and technical leadership as you go. A successful candidate has a strong software-engineering foundation, writes high-quality distributed-systems and services code, communicates clearly
  • is motivated to deliver results in a fast-paced, ambiguous environment. Key job responsibilities As a Software Development Engineer on the SageMaker AI team, you will: - Design, build, test
  • operate services that orchestrate foundation-model data preparation, training, evaluation
  • deployment as reliable, contract-validated workflows. - Own delivery of individual components end-to-end — from design and implementation through deployment, monitoring
  • on-call operations. - Build and extend compute-backend integrations and job launchers — submitting, monitoring
  • recovering large-scale training jobs across SageMaker (Training/Processing/HyperPod), EMR, AWS Batch
  • Kubernetes/EKS. - Improve the platform's resiliency and operability for long-running distributed jobs — checkpoint/resume, fault detection and recovery, retries
  • observability (metrics, logging, experiment tracking). - Contribute to the SDK, workflow orchestration
  • schema/contract layer that teams use to declare and run jobs
  • to the CDK infrastructure that deploys the platform. - Integrate containerized training and evaluation frameworks (e.g., PyTorch/FSDP, verl, NeMo/Megatron) into the platform's task and recipe model. - Contribute to design and architecture discussions, write clear technical designs
  • uphold engineering best practices (code review, testing, operational readiness). - Collaborate with ML scientists and internal customers to translate training requirements into reliable, self-service platform capabilities
  • help onboard and mentor interns and new engineers as you grow.
  • 3+ years of non-internship professional software development experience - 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience - 1+ years of software development engineer or related occupational experience - 1+ years of designing and developing large-scale, multi-tiered, multi-threaded, embedded or distributed software applications, tools, systems
  • services using: C#, C++, Java
  • Perl experience - 1+ years of Object Oriented Design experience - Bachelor's degree or foreign equivalent in Computer Science, Engineering, Mathematics
  • a related field - Experience programming with at least one software programming language
  • 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing
  • operations experience - Bachelor's degree in computer science or equivalent - - Experience with workflow/pipeline orchestration (Airflow, Step Functions
  • similar) and event-driven or service-oriented architectures. - - Experience with container and cluster compute (Kubernetes/EKS, Ray, Slurm, AWS Batch) and cloud infrastructure-as-code (AWS CDK/CloudFormation). - - Experience building and operating fully-managed cloud services at scale, including resiliency, checkpointing
  • fault tolerance for long-running jobs. - - Familiarity with machine-learning / deep-learning training workflows, GPU/accelerator compute (SageMaker HyperPod, AWS Trainium, P5-class GPUs)
  • distributed-training frameworks (PyTorch FSDP, Megatron-LM, DeepSpeed, verl). Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability
  • other legally protected status. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications
  • location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off
  • parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits . USA, WA, BELLEVUE - 143,700.00 - 194,400.00 USD annually

Required skills

GoJavaExpressAirflowPyTorchLLMAWSKubernetesDistributed Systems
Posted on JobRush — the end-to-end AI job-search platform.