ForgeApply · Job listing
Senior Site Reliability Engineer
Meridianlink
Apply in about a minute — without sacrificing quality.
ForgeApply autofills this application and tailors your resume to this exact posting. You review everything before it's sent. Free trial, no card required.
About this role
As a Senior Site Reliability Engineer on our cloud engineering team, you'll keep our production environment healthy, secure, and running smoothly. This is an operations-focused role: you'll own the day-to-day administration of our AWS accounts and databases, backup posture across our data stores, and production monitoring and debugging for a fully serverless platform. Your work will span the operational side of the software development life cycle — from deployment to maintenance and updates — always striving for continuous improvement. You'll keep our infrastructure clean, easily deployable, and scalable, creating a stable operating environment for the whole team.
Responsibilities
- Own day-to-day administration across AWS services, accounts, and access, as well as database administration across PostgreSQL and our other data stores.
- Own backup posture across databases, S3 buckets, and queues; verify restores regularly and maintain a tested disaster recovery plan.
- Proactively monitor production — CloudWatch dashboards, metric alarms, log-based metrics, and Slack alerting — addressing operational issues before they impact users.
- Lead production debugging and incident response: build and maintain runbooks, participate in the on-call rotation, and resolve queue and dead-letter-queue failures through retry, redrive, and recovery.
- Continuously refine our infrastructure to ensure it is easily deployable and scalable: keep infrastructure as code (SST/Pulumi) accurate, retire unused infrastructure, and keep cost visible and justified.
- Share your knowledge of production operations with the team, fostering a culture of learning and growth.
Qualifications: Knowledge, Skills, & Abilities
- Bachelor's degree and 4-6 years of related experience or equivalent work experience.
- 5+ years of experience in DevOps, site reliability, or platform operations, with significant responsibility for production systems.
- 3+ years of hands-on experience with AWS, with an emphasis on serverless services (Lambda, SQS, EventBridge, CloudWatch, S3).
- Strong database administration experience: PostgreSQL operations, backup and recovery, and query performance; comfort administering other data stores.
- Proficiency in scripting languages such as TypeScript, Python, and bash for production automation and operational tooling.
- Strong understanding of Linux, DNS, TLS, Docker, GitHub Actions, and infrastructure as code (SST, Pulumi, or Terraform).
- Experience with production monitoring and alerting, incident response, and on-call ownership.
Ready to apply to Meridianlink?
Apply in about a minuteSimilar jobs
- Senior Site Reliability Engineer — Epickids · Remote
- Senior Site Reliability Engineer — Drata · Hybrid - San Francisco
- Senior Site Reliability Engineer — Tulip · Somerville, MA
- Senior Site Reliability Engineer — Alembic · San Francisco HQ
- Senior Site Reliability Engineer — Airbyte · San Francisco
- Senior Site Reliability Engineer — Juullabs · United States of America
- Senior Site Reliability Engineer — Ixllearning · San Mateo, CA
- Senior Site Reliability Engineer — Ixllearning · Raleigh, NC
More like this: DevOps & SRE Jobs · Remote DevOps & SRE Jobs · More jobs at Meridianlink · Browse all jobs