Sr. System Development Engineer, AWS EC2 Manufacturing Infrastructure Services
Amazon Data Services, Inc.•3h ago
United StatesHybridFull-timeMid Level5+ yrs exp
H-1B verified · 2310 LCAs
Top focus
Infrastructure EngineerMl Infra EngineerAws EngineerSystem Admin
- Have you ever wondered what it would be like to build massively scalable systems that are used by the world's largest cloud infrastructures? Would you enjoy broad yet equally deep scope that impacts all AWS systems globally? AWS Manufacturing Infrastructure Services continues to pioneer and our team is architecting, building
- operating scalable services that are at the core of AWS infrastructure. What do we do? We own all AWS platforms, services, infrastructure
- tools that ensure the health of AWS hardware by testing every new system across all AWS manufacturing sites and ensure they are healthy when delivered to data centers. Our platform enables service owners such as EC2, EBS, S3
- other to deliver healthy servers for their service. Our team leads a large-scale service that sets the bar for Amazon and the industry in platform level services, effectively enabling the hardware at scale by designing and developing the software that manages the verification and testing of every server and rack in AWS manufacturing sites. Why it’s high-impact? We set the bar high to ensure that AWS customers get the capacity they need to run their applications on a healthy hardware server. What’s the challenge? There are many ambiguous and difficult challenges in our fast-moving space. Our platform is mission critical and requires deep system and software expertise. You need to have the ability to work within a fast moving and startup-like environment in a large company. You will identify solutions, trying ideas, given space to fail and iterate to produce products that your customers love. What you will do? You will be a part of a team to build the next generation of platform level software and systems that enables us to deliver healthy hardware to AWS customers. You design and deliver technology solutions which solve difficult business problems. Who would succeed in this role? Deeply technical engineers, who stay close to the customer as well as the systems architecture and design. They think about customer experience and the outcome. A person who works autonomously and dives deep in to a problem to deeply understand how things work, when to make subtle change
- when to disrupt the status quo to achieve the right results. Key job responsibilities - Own the reliability of manufacturing networking, systems
- platform infrastructure across AWS manufacturing partner sites globally. - Develop infrastructure-as-code to stand up and manage distributed manufacturing services at new and existing sites. - Build monitoring, alerting
- anomaly detection systems that identify infrastructure failures before they impact manufacturing throughput. - Develop troubleshooting tools and runbooks that enable rapid diagnosis and resolution of site infrastructure issues. - Consult with ODM/CM IT teams on network design and infrastructure standards, working across organizational boundaries where partners don't report to AWS. - Drive site expansion readiness, ensuring new manufacturing sites are infrastructure-ready within target onboarding timelines. - Implement security and operational best practices for manufacturing environments, including credential management, network segmentation
- access controls. - Contribute to architecture decisions for manufacturing platform evolution. - Develop and build systems that enables operating manufacturing infrastrcure operation at scale About the team We are EC2 Manufacturing Infrastructure Services. We exist to make sure that a working rack delivers to our customers in AWS data centers. Every server that powers EC2, every rack that runs AI/ML workloads, every piece of capacity that AWS customers depend on passes through our systems before it ships. We own Server Level Testing, Rack Level Testing, Firmware Upgrade services, code deployment into manufacturing sites, network architecture at contract manufacturers
- all infrastructure installed at those sites. When yield drops or a test pipeline breaks, we are the team that responds. Nothing reaches an AWS data center without passing through us first. If our services go down, manufacturing lines stop. If our infrastructure is unreliable, capacity doesn't ship. The blast radius of what we do is measured in billions of dollars of customer workloads. We are a small team with enormous scope. You will own real systems on day one. We don't have layers of abstraction between you and the problem. You build it, you ship it, you operate it, you improve it.
- 3+ years of programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby experience - 5+ years of non-internship professional software development experience - 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience - Deep expertise in Linux OS and network troubleshooting, including managing full application stacks from the OS up through custom applications in distributed, multi-site environments - Experience leading enterprise scale infrastructure or development-based cloud programs/projects, including defining technical direction and driving adoption across teams - Experience designing and implementing automation frameworks at scale using Python, Java, Perl, PHP, Ruby, Bash, Shell
- equivalent - 5+ years of systems engineering, network engineering
- SRE experience, including ownership of reliability strategy for distributed infrastructure - Advanced networking: TCP/IP, DNS, DHCP, VLANs, routing, firewalls, network segmentation, including network architecture design for multi-site environments - Deep knowledge of systems engineering fundamentals (networking, storage, operating systems, firewalls) with experience setting standards for distributed environments - Experience mentoring engineers and driving technical decisions across team boundaries
- Experience working in an Agile environment using the Scrum methodology - 3+ years experience in data centers, infrastructure service providers, or manufacturing technology environments - Experience designing and implementing security architectures (Authentication, Authorization, SSO, Cryptography, credential rotation, zero-trust principles) - Deep experience with AWS services in edge/hybrid environments (EC2, VPC, IAM, Outposts, CloudWatch, S3), including architecture ownership - Experience architecting on-premise infrastructure solutions (servers, storage, campus networking, firewalls) across multiple sites - Hands-on experience with PXE boot, IPMI/BMC, or bare-metal provisioning at scale - Experience designing monitoring and observability architectures (CloudWatch, Prometheus/Grafana, or similar) including anomaly detection and automated remediation - Experience leading technical relationships with external partners or vendors, including defining standards and driving compliance - Experience with manufacturing environments, industrial networking, or factory automation systems, including driving reliability improvements across sites - Experience owning on-premise IT environment strategy and driving improvements across organizational boundaries - Track record of identifying systemic issues and delivering solutions that eliminate entire classes of problems Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff
- adhere to standards of excellence despite stressful conditions
- communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service
- and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits . USA, CA, Cupertino - 173,900.00 - 235,200.00 USD annually
Required skills
C++C#JavaPythonGolangPowerShellRubyLinuxnetwork troubleshootingautomation frameworksTCP/IPDNSDHCPVLANsfirewalls