ForgeApply · Job listing
Staff Software Engineer - AIOps
Earlywarning
See all 77 open roles at Earlywarning →
Tailor your resume for this Earlywarning job in about a minute.
ForgeApply tailors your resume and cover letter to this exact posting, then hands you a ready-to-submit application for Earlywarning's site. Free trial, no card required.
About this role
At Early Warning, we’ve powered and protected the U.S. financial system for over thirty years with cutting-edge solutions like Zelle®, Paze℠, and so much more. As a trusted name in payments, we partner with thousands of institutions to increase access to financial services and protect transactions for hundreds of millions of consumers and small businesses.
Positions located in Scottsdale, San Francisco, Chicago, or New York follow a hybrid work model to allow for a more collaborative working environment.
Candidates responding to this posting must independently possess the eligibility to work in the United States, for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.
Staff Software Engineer - AIOps Overall Purpose The Staff Software Engineer - AIOps is a senior hands-on technical contributor responsible for designing, building, and operating the AIOps agent platform that enables AI agents to observe events across Early Warning technology environments, generate diagnoses and recommendations, and execute approved actions within defined safety controls. The role focuses on the software systems, control loop, tool interfaces, telemetry, evaluation capabilities, and operator experience required to run agentic capabilities safely and reliably in production; it is not primarily a model development or research role. This position executes complex technical work with general direction, independently makes implementation decisions within established architecture and standards, and may seek guidance for novel, highly ambiguous, or cross-enterprise decisions. The Staff Engineer applies software engineering, distributed systems, and platform engineering principles and partners across Engineering, Technology Operations, Security, Risk and Compliance, and other technology teams to establish reusable patterns and guardrails for progressive levels of automation. Essential Functions • Apply mature software engineering practices across the agent runtime, tool layer, APIs, and operator experience, including versioned interfaces, automated testing, code review, release management, observability, and regression coverage. • Execute complex AIOps engineering assignments with general direction; break work into deliverable components, identify practical implementation solutions, raise risks and dependencies, and seek guidance when decisions extend beyond established architecture or standards. • Contribute to the design and evolution of the agent platform architecture; maintain architecture decision records, define interface standards, and evaluate significant design and build-versus-buy tradeoffs with appropriate technical guidance. • Design, build, test, and operate the core agent runtime, including event ingestion, context assembly, planning, tool invocation, verification, escalation, and response handling. • Design and maintain scoped, typed, and auditable tool interfaces that enable agents to interact with CI/CD platforms, infrastructure automation, Kubernetes, network services, ITSM platforms, observability systems, and secrets management solutions. • Define and implement controls for agent actions, including read-only and advisory capabilities, human approval requirements, narrowly scoped unattended actions, permission boundaries, blast-radius limits, dry-run capabilities, reversibility, and emergency shutdown controls. • Define tool contracts, versioning standards, testing requirements, and intended behavior so agent-accessible capabilities are managed as reliable production interfaces. • Build event-routing capabilities that receive and prioritize alerts, pipeline failures, tickets, operational requests, and other technology events and provide appropriate context to the agent runtime. • Establish telemetry capabilities used by the agent, including logs, metrics, and traces from relevant technology platforms, and contribute to standards for telemetry quality and schema evolution. • Design and operate model routing, session and state management, retries, timeouts, and cost and latency controls appropriate for a production service. • Implement comprehensive auditability for events received, decisions generated, approvals obtained, and actions executed to support operational, risk, compliance, and examination requirements. • Develop evaluation capabilities for agent decision quality, including historical incident replay, controlled or shadow-mode evaluation, measurement of proposed actions, and regression testing. • Build feedback mechanisms that incorporate human approvals, overrides, and operational outcomes into measurable improvements to platform quality. • Package reusable agent capabilities, tool integrations, APIs, documentation, and implementation patterns so other technology teams can extend and adopt the platform through defined self-service practices. • Develop and maintain the operator experience, including approval workflows, agent activity and decision history, telemetry views, and APIs that expose agent state to user interfaces and other systems. • Partner with Engineering, Technology Operations, Architecture, Security, Risk, and Compliance stakeholders to evaluate technical tradeoffs and ensure solutions meet reliability, security, operational, and regulatory requirements. • Define and monitor measures of platform effectiveness, including recommendation quality, approval and override rates, action success, rollback frequency, latency, cost, adoption, and operational outcomes. • Support the company's commitment to risk management and protecting the integrity and confidentiality of systems and data.
Minimum Qualifications • Education and/or experience typically obtained through a bachelor's degree in computer science, engineering, or a related technical field. • Typically eight or more years of related experience in software engineering, platform engineering, distributed systems, DevOps, AIOps, or a related technical discipline. • Demonstrated str
Salary insight
The midpoint of this range ($145k) is about 11% above the median disclosed salary for Chicago roles listed on ForgeApply ($130k across 2,704 jobs).
See full Software Engineer salary data for Chicago →
Based on live postings with disclosed pay on ForgeApply; refreshed daily. Not an estimate of this employer's offer.
Tailor your resume for this Earlywarning role before you apply.
Tailor my resume for this jobSimilar jobs
- Staff Software Engineer - AI — Cloudera · US-California-San Jose
- Staff Software Engineer (AI CICD) — Chainguard · Remote
- Staff Software Engineer (AI) — Pendo · New York, NY
- Staff Software Engineer- AI Workload Orchestration — Coreweave · Sunnyvale, CA / Bellevue, WA
- Staff Software Engineer - Infrastructure Automation — Earlywarning · Scottsdale | Chicago
- Staff Software Engineer, AI Engineering — Tebra · Remote
- Staff Software Engineer I — Confluent · Remote
- Staff Software Engineer I — Thomson Reuters · Minnesota, Eagan, United States | Texas, Frisco, United States | New York, New York, United States
More like this: Software Engineer Jobs · Software Engineer Jobs in Chicago · Browse all jobs
Free ATS checker · How to Tailor Your Resume to a Job Description (Step by Step)