ForgeApply · Job listing
Staff Production Operations Engineer
Sony Interactive Entertainment
See all 114 open roles at Sony Interactive Entertainment →
Tailor your resume for this Sony Interactive Entertainment job in about a minute.
ForgeApply rewrites your resume for this exact posting, then autofills the application on Sony Interactive Entertainment's site with it. You review everything before it's sent. Free trial, no card required.
About this role
Why Sony Interactive Entertainment?
Sony Interactive Entertainment isn’t just the Best Place to Play — it’s also the Best Place to Work. Sony Interactive Entertainment (SIE) is the company behind the PlayStation brand. As a subsidiary of Sony Group Corporation, we’re part of a proud legacy of innovation and excellence. SIE is a dynamic technology company, delivering cutting-edge hardware and network services to more than 100 million people and an entertainment leader, home to some of the most beloved and recognizable intellectual properties (IP) in the world. Our role at SIE is to create and nurture the experiences under the PlayStation brand, a name synonymous with entertainment excellence and creativity.
As a member of the Production Operations Engineering team within the platform technology group, you will help keep key user experiences on the platform available, resilient and high performing across time zones and critical business periods, while continually enabling our service teams to deliver new and exciting products and technical features. The team is trusted to respond when production services need support, including through rotational on-call, incident response and urgent operational escalations. You will be empowered to define and lead technical initiatives, helping identify and proactively drive improvements in process, technology and production operations supporting millions of users.
This Staff individual contributor role carries broad technical ownership: a recognized expert and hands-on system architect whose decisions shape reliable production services across multiple partner teams. The ideal candidate brings deep full-stack engineering and production experience, strong automation instincts, cloud/container operations expertise, incident leadership, and curiosity for using AI-assisted workflows to improve operational quality, release confidence and delivery speed.
Responsibilities:
• Lead application operations and production support for internal and public-facing services in AWS and container environments, with focus on availability, resiliency, scalability, performance and security.
• Define and evolve production operations architecture, readiness standards and reusable patterns for services/features, including automation, monitoring, runbooks, rollback plans and operational acceptance criteria.
• Architect and scale tools, scripts and automation frameworks across API/service telemetry, release validation and operational readiness to reduce toil and improve incident response.
• Establish CI/CD operational gates, production validation practices and governance mechanisms that improve release confidence, scalability and engineering velocity.
• Advance observability, event correlation and operational metrics using service health signals, automation health, release risk indicators, dashboards and alerts that reduce MTTD and MTTR.
• Partner with SRE, data services, CI/CD, service engineering, platform hosting, product and business teams to shift reliability left and improve shared ownership of production quality.
• Lead performance, capacity, infrastructure modernization and cost optimization for services using AWS, Kubernetes, EKS/UKS, autoscaling and related cloud patterns.
• Use AI-assisted and agentic tools, such as Codex, Claude, Cursor or similar, to accelerate development, operational automation and production-signal analysis.
• Provide rotational on-call support and improve incident workflows using Slack, JIRA, ServiceNow, BigPanda and related escalation systems; lead RCA and systemic fixes for recurring production issues.
Key Qualifications:
• Staff-level hands-on engineer equally comfortable with software development, systems engineering, production operations and service architecture.
• System Architect and full-stack engineer who connects customer experience, application behavior, APIs, data flows, infrastructure and operational controls into resilient designs.
• Proven ability to define operational standards, quality gates, reusable frameworks and governance practices across teams.
• Independent technical leader who influences cross-functional partners, mentors engineers and drives adoption of scalable practices.
• Strong systems-thinking and troubleshooting depth across runtime, cloud, Kubernetes/EKS/UKS, network, observability, release and incident layers.
Required Skills:
• Distributed service architecture and operations, Unix/Linux systems internals, networking, TCP/IP, HTTP/HTTPS, DNS, load balancing and API troubleshooting.
• AWS service operations, including ALB, Route 53, API Gateway, Lambda, RDS, DynamoDB, ElastiCache and Java/API services.
• Containers and orchestration, including Docker, Kubernetes, EKS/UKS and Fargate.
• Automation and software development in Python, Go or Java; source control and configuration management using GitHub/Git, Ansible, Chef or similar.
• Infrastructure as Code and CI/CD using Terraform, CloudFormation, Jenkins, Spinnaker or comparable tooling.
• Observability, incident management and collaboration tooling such as Datadog, CloudWatch, Splunk, Grafana, BigPanda, Slack, JIRA and ServiceNow.
• Strong communication skills with the ability to adjust messaging by audience, translate complex technical issues for cross-functional business partners, and articulate both technical detail and customer/business impact.
• Data reporting, analytics or operational data platforms such as SQL, MySQL, Oracle, Snowflake or big data systems.
• AI-assisted development or agentic engineering workflow experience using tools such as Codex, Claude, Cursor or similar.
Experience:
• BS degree or equivalent in Computer Science, Software Engineering or related technical area.
• 10+ years operating and supporting services in production environments at scale.
• 5+ years AWS Cloud experience deploying, tuning and operating Java/API services.
• Hands-on incident management experience, including on-call, incident coordinat
Salary insight
This posting doesn't disclose pay. Across 804 San Diego jobs with disclosed salaries on ForgeApply, the median is $147k.
See full Operations salary data for San Diego →
Based on live postings with disclosed pay on ForgeApply; refreshed daily. Not an estimate of this employer's offer.
Tailor your resume for this Sony Interactive Entertainment role before you apply.
Tailor my resume for this jobSimilar jobs
- Staff Production Engineer — Zscaler · Remote - California, USA; San Jose, California, USA
- Staff Production Engineer — Northrop Grumman · United States-Minnesota-Plymouth
- Senior Production Operations Engineer — Sonyinteractiveentertainmentglobal · United States, San Diego, CA
- Production Operations Engineer — Freeformfuturecorp · Los Angeles, CA (On-site)
- Staff Network Production Engineer, Operations — Crusoe · San Francisco, CA - US
- Staff DevOps Engineer — Generac · Denver, CO - USA | Waukesha, WI - USA
- Staff DevOps Engineer — Northrop Grumman · United States-Colorado-Aurora | United States-Virginia-Fairfax
- Staff DevOps Engineer — Archer56 · San Jose, California, United States
More like this: Operations Jobs · Operations Jobs in San Diego · Browse all jobs
Free ATS checker · How to Autofill Greenhouse Job Applications (Without Sending Junk)