Site Reliability Engineer I
VeritoneAbout the role
WE ARE VERITONE (
We are driven by the belief that Artificial Intelligence is mankind’s greatest invention. It is the key to building a safer, more vibrant, transparent, and empowered society. We are determined to be an active contributor to shaping our future for the better. We care about the ethical implications of AI and the prosperity and well-being of all individuals, as well as the growth and continued successes of our employees, customers, and partners.
Veritone’s mission today is more important than ever. We’re here to democratize AI and enable every organization and every person with the power of AI. What started in 2014 with the idea of providing unified access to hundreds of cognitive engines through one common software infrastructure, evolved to the world’s first AI operating system, aiWARE, which orchestrates a diverse ecosystem of cognitive engines to power intelligent automation for both commercial and government organizations. As we progress, we will continue to move humans from “in” to “on” to “out of the loop” to help them accelerate workflows, save time and costs, and uncover new insights and opportunities. You can view us at: www.veritone.com / www.veritoneone.com
What You'll Do
- Deploy and maintain a resilient, secure, and efficient SaaS application platform to meet established SLAs.
- Automate, monitoring, management and incident response to achieve an auto-remediation system.
- Monitor site stability and performance and troubleshoot site issues.
- Scale infrastructure to meet rapidly increasing demand.
- Manage cross-functional requirements working with Engineering, Product, Services, and other departments.
- Collaborate with developers to bring new features and services into production.
- Independently design and develop tools to aid in operations and automation as well as work jointly with other team members to deliver innovative solutions to complex business and technical challenges.
- Provide deployment and operations support for multi-tiered distributed software applications.
- Estimate engineering effort, plan implementation, and rollout system changes that meet requirements for functionality, performance, scalability, reliability, and adherence to development goals and principles.
- Collaborate in a fast paced environment with multiple teams (software development, release management, build and release, etc...).
- Collaborate in a fast paced environment with multiple teams in a dynamic entrepreneurial organization
- Defining how the behavior of large scale systems can be achieved
- Measuring and achieving reliability through engineering and operations work
- Monitoring and alert development, documentation and management with the goal of creating an auto-remediation system
- Adapting security controls to product not typically native to GA releases
- Developing automation methods to extend standard deployment pipelines for bespoke implementations
- Patching, policy enforcement, and audit of production systems
- Driving the Disaster Recovery process
What You'll Need
- Expertise with Terraform and/or Ansible.
- Knowledge of JavaScript, Go, or other programming languages
- 3+ years of professional Linux systems and software management experience
- Expertise with Infrastructure-as-Code including Ansible and Terraform
- Knowledgeable with code languages including: Go, Node.js, Java
- Experience with managing infrastructure within Azure, GCP and AWS
- Expertise with monitoring and alerting systems including Prometheus, Grafana
- Strong script skills for systems and data driven solutions
- JIRA experience for project/task management
- Extensive experience in troubleshooting large-scale distributed systems.
- Strong background working in AWS, GCP, Azure and general Linux environments.
- Comprehensive background in monitoring and alerting systems in auto-remediation systems.
- Proven examples of standardizing security controls across large-scale systems
- Comfort working within project/task management platforms
Systems and Tools
- Cloud platforms including: Azure, GCP and AWS
- Infrastructure coding languages: Terraform, Cloudformation, Ansible, Puppet
- CI/CD: experience working with and supporting build and deploy pipelines and tools: Jenkins, GitHub Actions, Rundeck
- Datastore Management and Query skills: Postgres, MySQL, Mongo, ElasticSearch, Solr
- Container orchestration platforms: Docker, Kubernetes, EKS, AKS
- Familiarity with coding languages including: Go, Node.js,
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s