ForgeApply
Try it free

ForgeApply · Job listing

Principal Software Engineer - Observability & Telemetry Data

Smartsheet

Bellevue, WA, USonsite

See all 59 open roles at Smartsheet

Tailor your resume for this Smartsheet job in about a minute.

ForgeApply rewrites your resume for this exact posting, then autofills the application on Smartsheet's site with it. You review everything before it's sent. Free trial, no card required.

About this role

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day.

Observability Engineering owns the breadth and depth of observability and telemetry at Smartsheet. We build and operate the shared platform every engineering team depends on to understand how their software behaves in production (metrics, logs, distributed traces, RUM, synthetics, and AI/agentic telemetry) across a global, multi-region, commercial and government footprint. We are consolidating ownership of all of it into a single correlated platform, and treating telemetry as what it actually is: one of the largest and most valuable datasets at the company. It's a broad mandate with real scope to shape, and we hold this internal platform to the same rigor, reliability, and product standards as customer-facing software.

As a Principal Software Engineer (Observability & Telemetry Data) , you will set the technical direction for how Smartsheet collects, models, stores, and derives value from telemetry data. This is an architect role: you will define the telemetry data platform and the standards that govern it, carry that architecture across engineering org boundaries, partner with our Data and AI Platform teams on Databricks and MLflow integration, and act as the company's definitive technical voice on observability. You will influence roadmaps you do not own, resolve the hard design trade-offs between fidelity and cost, and raise the technical bar for every engineer instrumenting a service.

This full-time position reports to the Team Lead, Observability Engineering and can be located in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer.

You Will

• Architect the Telemetry Data Platform: Own the end-to-end design for how telemetry is landed, modeled, and queried, including open table formats, partitioning and schema evolution strategy, separation of storage and compute, tiered retention, and the query interfaces engineers actually use. Make telemetry a durable, portable, Smartsheet-owned dataset rather than a vendor-locked byproduct.

• Own the Telemetry Data Model: Define the semantic conventions, shared business identifiers (user, org, plan, tenant), and schema standards that let any signal be correlated with any other, and drive their adoption across every service team at Smartsheet.

• Set OpenTelemetry Direction: Lead the migration to OTel-based instrumentation, defining collector architecture, context propagation, and sampling strategy (including tail-based sampling) so that engineers can move from a log line to a trace to a metric without losing the thread.

• Integrate AI and Agentic Telemetry with Databricks: Own the architecture connecting our observability platform to Databricks and MLflow, so that agentic and model telemetry (prompt, completion, tool and MCP calls, evaluation results) is captured un-sampled, stays useful under the input/output redaction our governance requires, and reconciles cleanly with the traces and metrics in our primary observability stack.

• Instrument the Data Platform Itself: Bring first-class observability to our data estate, including Databricks jobs, pipelines, and warehouses, with meaningful signals for freshness, data quality, lineage, and cost, so data reliability is measured with the same discipline as service reliability.

• Build the Analytics Layer on Telemetry: Turn telemetry into decision-grade analytics, covering reliability and incident metrics, telemetry cost and chargeback models, and adoption and coverage reporting that leadership can act on.

• Engineer Collection and Routing at Scale: Architect the high-volume collection and routing tier (FluentBit, Kinesis, and OTel collectors) that moves telemetry from every service to its destination across US, EU, AU, and GovCloud regions, and own the migration of these pipelines as we consolidate onto a unified backend.

• Own Telemetry Economics: Set the cost architecture for observability data, including ingest governance, cardinality control, and storage tiering, so that teams get the fidelity they need to debug without the spend pressure that causes them to under-instrument.

• Carry Architecture Across Org Boundaries: Partner with the Data Platform, AI Platform, and infrastructure organizations to align telemetry architecture with theirs, influence roadmaps you do not own, and represent observability in company-level platform and vendor decisions.

• Raise the Technical Bar: Lead design and code reviews, author the architecture decisions and standards others build against, and mentor senior and mid-level engineers on instrumentation, telemetry data modeling, and cost-aware design.

• Participate in a production support and on-call rotation, taking ownership of the most complex issue resolution and driving root-cause analysis that improves system resiliency.

You Have

• 10+ years of experience building and operating large-scale distributed systems, data platforms, or observability infrastructure, including time at Principal or Staff level.

• Deep Observability Expertise: Hands-on production ownership of metrics, logs, and distributed tracing at scale, including at least one major backend (Datadog, or comparable) and a working understanding of cardinality and cost mechanics.

• Telemetry Data Engineering: Demonstrated depth in large-scale data architecture, including open table formats (Delta Lake, Apache Iceberg), Spark or comparable distributed processing, streaming ingestion, partitioning and schema evolution, and query performance and cost tuning over very large datasets.

Salary insight

This posting doesn't disclose pay. Across 1,467 Seattle jobs with disclosed salaries on ForgeApply, the median is $158k.

See full Software Engineer salary data for Seattle

Based on live postings with disclosed pay on ForgeApply; refreshed daily. Not an estimate of this employer's offer.

Tailor your resume for this Smartsheet role before you apply.

Tailor my resume for this job

Similar jobs

More like this: Software Engineer Jobs · Software Engineer Jobs in Seattle · Browse all jobs

Free ATS checker · How to Autofill Greenhouse Job Applications (Without Sending Junk)