ForgeApply · Job listing
GenAI Platform Engineering Lead
M&T Bank
See all 113 open roles at M&T Bank →
Tailor your resume for this M&T Bank job in about a minute.
ForgeApply tailors your resume and cover letter to this exact posting, then hands you a ready-to-submit application for M&T Bank's site. Free trial, no card required.
About this role
Manages the activities of Engineering Team Leaders, Engineering Supervisors and/or engineering units responsible for the reliability, observability and operational readiness of the Bank’s AI platform. Provides day-to-day direction for the teams and applications in alignment with departmental goals and the needs of the clients they support. Serves as the technical lead for AI operations and service assurance, with accountability for the operational layer of the AI platform. Works closely with the Platform Engineering Manager, who owns the broader platform, to ensure production systems are secure, resilient, observable, cost-effective and operationally ready. Oversees incident response practices, runbooks, monitoring, infrastructure automation and production support. Responsible for managing client relationships and expectations, prioritizing the project queue and achieving individual and organizational objectives at minimum cost. Primary Responsibilities • Lead the reliability, observability and operational readiness of the AI platform, including production monitoring, incident response, service assurance, infrastructure automation and operational controls. • Establish and maintain operational runbooks, escalation procedures, service-level indicators, service-level objectives and incident response practices. • Measure operational performance through platform availability, incident response time, mean time to recovery, model latency, token consumption and infrastructure cost. • Provide leadership during production incidents. Coordinate troubleshooting, communication, escalation, root-cause analysis and corrective actions. • Build and maintain observability pipelines using OpenTelemetry, Prometheus, Grafana, Azure Monitor, Log Analytics and/or comparable technologies. • Oversee the deployment and operation of Azure infrastructure and services, including Azure API Management, Azure Monitor, Log Analytics, Microsoft Entra ID and infrastructure managed through Terraform. • Promote infrastructure-as-code and automated deployment practices using Terraform, GitHub Actions, GitLab CI and/or comparable tools. Ensure platform changes are deployed through controlled pipelines rather than manual processes. • Drive automation of recurring operational activities using Python, Bash, PowerShell and other appropriate scripting technologies. • Oversee AI-specific operational capabilities, which may include token cost tracking, model latency monitoring, provider failover, caller-level rate limiting, prompt logging, personally identifiable information interception and infrastructure-layer content filtering. • Partner with cybersecurity and risk teams to support Security Information and Event Management integrations and security event feeds from application infrastructure. • Manage and participate in consultations with client management to analyze short-range business requirements and recommend innovations that anticipate the future impact of changing business and technology needs. Build and maintain positive client relationships. • Monitor technology direction, industry trends and vendor applications related to AI platforms, site reliability engineering, cloud infrastructure, observability and service assurance. • Research and initiate changes to existing processes, technologies and operating models when necessary. • Lead vendor and product analysis and provide recommendations. • Oversee application development support, testing efforts, technology infrastructure, project management and other assigned technology domains. • Serve as a subject matter expert for AI platform operations, production reliability, observability and service assurance. • Build rapport across the organization and maintain a professional level of communication and cooperation with technology, business, risk, cybersecurity and vendor partners. • Maintain relationships with vendors and professional organizations. • Direct team activities, assign personnel to projects and provide technical and operational guidance. • Ensure schedules and commitments are completed. Lead short-term staffing and capacity planning. • Implement technology consistent with Division standards and long-range plans. Ensure adherence to Department and Technology standards and procedures, including documentation, audit trail and change management requirements. • Translate business and operational requirements as needed to assist staff in preparing detailed specifications for system enhancements. • Evaluate and manage recommended designs based on business, operational and technology requirements. Identify, communicate and resolve issues and concerns. • Manage project plans and coordinate major project and production-readiness activities. Remain current on work outside the team that may affect the team, platform or client environment. • Develop and manage multiple cost center budgets, including cloud infrastructure and AI platform operating costs. • Recommend and implement policies and procedures that improve the performance, reliability and effectiveness of the Department. • Exercise the usual authority of a manager concerning staffing, performance appraisals, promotions, salary recommendations, performance management and terminations. • Understand and adhere to the Company’s risk and regulatory standards, policies and controls in accordance with the Company’s Risk Appetite. Design, implement, maintain and enhance internal controls to mitigate risk on an ongoing basis. Identify risk-related issues requiring escalation to management. • Promote an environment that supports belonging and reflects the M&T Bank brand. • Maintain M&T internal control standards, including the timely implementation of internal and external audit points and resolution of issues raised by external regulators, as applicable. • Complete other related duties as assigned.
Scope of Responsibilities Oversees a team where the majority of employees are engineers, architect individual contributors, Engine
Salary insight
The midpoint of this range ($155k) is about 35% above the median disclosed salary for Buffalo roles listed on ForgeApply ($115k across 204 jobs).
See full DevOps / SRE salary data for Buffalo →
Based on live postings with disclosed pay on ForgeApply; refreshed daily. Not an estimate of this employer's offer.
Tailor your resume for this M&T Bank role before you apply.
Tailor my resume for this jobSimilar jobs
- GenAI Platform Engineering Manager — M&T Bank · Buffalo, NY
- GenAI Engineer — Accenturefederalservices · Washington, DC
- GenAI Senior Software Engineer — Visa · US - Austin, TX | US - Bellevue, WA
- Senior Lead AI Engineer (GenAI Platform Services) — Capital One · New York, NY | San Francisco, CA | San Jose, CA
- Head of Engineering, AI Platform — Gusto · San Francisco, CA - Hybrid
- Head of AI Platform Engineering — Postman · San Francisco, California, United States
- Lead AI Platform Engineer — OUTFRONT · New York, NY | Fairfield, NJ
- Principal Engineer - AI Platform — Iherb · Home Office, CA
More like this: DevOps & SRE Jobs · DevOps & SRE Jobs in Buffalo · Browse all jobs
Free ATS checker · How to Tailor Your Resume to a Job Description (Step by Step)