Jobs and Careers
AR

Senior Manager - IT Operation Strategic with Development Experience

Arthur Grand Technologies Inc
Alexandria, United Statesfull_timeVerifiedPosted 21 Dec 2023

About the role

Company Description

Arthur Grand (AG) is an IT services firm specializing in Digital Transformation initiatives for Federal, Commercial, State & local customers. Since 2012, AG has (been) successfully supporting and delivering IT services to our customers in the areas of enterprise modernization and transformation with a core focus on emerging technologies including Cloud Solutions (AWS, Azure), Agile Development and Custom Programming, Full Stack Development, DevOps, DevSecOps, & CI/CD, Web & Mobile APP Development, Data Visualization and Data Warehousing, Financial/ERP System Implementation, Infrastructure Management. Arthur Grand’s culture of (delivery excellence) or excellent delivery, combined with a commitment to bring the best talent to provide services, has earned our company an unparalleled reputation for delivering transformative results.

Job Description

Position: IT Operations Strategist (Senior Manager & Director)

Location: Remote and possibly on-site pending return to work , if onsite is required in the future, location is expected to be in the Alexandria, VA area.

Duration: Contract to Hire

 

You will play a critical role in solving impactful operational problems for our clients. You will think creatively to find opportunities to improve observability, system performance and efficiency, scalability, fault tolerance, and self-healing capabilities. You’ll apply Resiliency and Chaos Engineering principles to challenge the client’s systems and discover hidden weaknesses, all while understanding the big picture of how systems work together to create the ultimate client experience.

 

Minimum requirements:

  • Bachelor’s degree, Minimum of eight years related work experience, with at least three years of development experience.
  • Undergraduate degree or equivalent combination of training and experience. Graduate degree preferred.
  • Full stack development in open source.
  • Ability to diagnose and resolve problems in high-throughput applications,
  • Experience with one or more observability frameworks or tools – Experience with Cloudwatch, Instana, Splunk, Harness Chaos Engineering, etc.
  • Strong understanding of database principles and working knowledge in distributed storage and infrastructural solutions.
  • Strong understanding of network principles and working knowledge in load balancing and network security.
  • Experience with container management and micro-services architectures in cloud and on-premises infrastructure.
  • Working knowledge of AWS network foundations, application networking, edge, and network security.
  • Excellent communication, and documentation skills.

 

Core Responsibilities:

  • Advises on instrumenting, enhancing, and advocating for system observability. Identifies and recommends solutions to bridge system observability gaps.
  • Collaborates with internal teams to evaluate the health, stability, and reliability of systems/platforms. Looks for opportunity to improve system availability, performance efficiency and resiliency.
  • Develops and communicates new standards and newly available tools and frameworks across teams. Includes automated solutions for reliability.
  • Provides technical leadership, consultancy, and coaching on designing and implementing both traditional and serverless architectures in AWS with an emphasis on repeatability, scaling options, resilience, reliability, telemetry, networking, etc., including design patterns for resilient systems
  • Leads failure modes analysis spanning product families when new features and architecture patterns are introduced. Leads cross-product or cross-subdivision chaos experimentation. Facilitates post-incident reviews for any high severity client impacting events local to the product family.
  • Designs, reviews, and coaches others on performance tests using appropriate components (e.g., requests per minute, # of threads, the construction of a request with headers and cookies)
  • Consults, reviews, coaches, and influences architectural decisions, including non-functional aspects, proposing potential technical solutions/enhancements, and explaining convincingly which is better and why.
  • Contributes to and supports Reliability Engineering and Resilience communities of practice. Remains informed about site reliability engineering activities happening within the organization.
  • Provides technical leadership, guidance, consulting, training, and governance on SRE to one or more product families. Works with product owners and teams to set goals for higher availability and SRE impact, and tracks progress toward achieving them.
  • Identifies

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Arthur Grand Technologies Inc

View company profile →