Jobs and Careers
SE

Senior Software Engineering Manager - FinOps Platform Services

ServiceNow
Pleasanton, United StatesRemotefull_timeVerifiedPosted 7 Aug 2026

About the role

Company Description

It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.

Join us to put AI to work for people.

 

Job Description

What you get to do in this role:

Platform Ownership & Operations 

  • Own the operational health and reliability of Trino, Lightdash, Coder, Jupyter, Redash, Hive Metastore, and Nessie across development and production environments. 

  • Establish and maintain SLOs for platform availability, query performance, and workspace provisioning. Build the dashboards and alerting to track them. 

  • Own Trino cluster operations end to end, including deployment, scaling, upgrades, performance tuning, resource group management, query optimization support, and user access controls. 

  • Drive the platform upgrade and patching cadence, balancing stability with staying current on security fixes and feature releases across all services. 

  • Build runbooks, on-call processes, and incident-response practices so the team can respond to and resolve production issues quickly and learn from them. 

  • Ensure platform security across all services, including access controls, authentication (SSO/OIDC integration), secrets management, and audit logging. 

Platform Evolution & Roadmap 

  • Lead the migration from Hive Metastore to Nessie as the versioned Iceberg catalog, delivering Git-like branching semantics, safe multi-writer coordination, and auditable catalog history. 

  • Drive Lightdash platform improvements including version upgrades, performance optimization, row-level security configuration, and the governed self-service analytics experience. 

  • Evolve the Coder platform through workspace template lifecycle management, resource policies, idle-stop tuning, and onboarding new users and use cases including AI coding agents. 

  • Own the Jupyter and Redash platforms, ensuring availability, scaling, integration with Trino and the lakehouse, and user lifecycle management. 

  • Evaluate and adopt new open-source technologies where they raise the platform’s ceiling or reduce operational burden. 

People Leadership 

  • Manage, mentor, and grow a team of 3 to 5 platform engineers. Set clear expectations, provide regular feedback, and create career development paths. 

  • Hire and build the team to match the platform’s growing scope and user base. 

  • Foster a culture of operational excellence, automation over toil, and blameless incident retrospectives. 

  • Set engineering standards for how the team builds, deploys, monitors, and documents platform services. 

Collaboration & Stakeholder Management 

  • Partner with the DevOps/infrastructure team on Kubernetes capacity, networking, storage, and CI/CD pipeline needs for your platform services. 

  • Serve as the platform liaison to data engineers, analysts, and FinOps practitioners. Understand their workflows, gather feedback, and prioritize improvements that unblock them. 

  • Collaborate with the Data Platform and Data Governance teams to ensure platform services align with enterprise standards for security, lineage, and access control. 

  • Support the broader Cloudera-to-lakehouse migration by ensuring Trino, Nessie, and the catalog layer are production-ready for migrated workloads. 

  • Apply AI/ML tooling where it accelerates platform operations, monitoring, or user support. 

What success looks like 

  • Platform services meet their SLOs consistently, and the team has the observability and processes to detect and resolve issues before users are affected. 

  • Trino queries perform reliably at scale with well-managed resource groups and a clear upgrade cadence. 

  • The Hive Metastore to Nessie migration is planned, sequenced, and executing without disruption to downstream users. 

  • Lightdash and Coder are stable, current, and adopted broadly across the organization with minima

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

ServiceNow

View company profile →