Jobs and Careers
AR

Staff Data Engineer (TS/SCI with Full Scope Polygraph) {S}

ARKA Group, LP
United Statesfull_timeVerifiedPosted 26 Feb 2026

About the role

ARKA Group L.P. (“ARKA”) is an advanced technologies company serving the U.S. military, intelligence community, and commercial space industry delivering next-generation solutions to support the national security space enterprise. Built on more than six decades of excellence, ARKA brings modern approaches and a culture of innovation to the challenges of today.

Join the ARKA team to learn how Beyond Begins Here. Discover your next career opportunity now!

Position Overview:

We are looking for a Staff Data Engineer/Scientist looking for new challenging problems. You will support the development of AI/ML algorithms in a multitude of disciplines from large language models, natural language processing, and time-series predictive analytics. Additionally, we have a team of excellent researchers and software developers who are eager to mentor and teach their craft.  We have multiple positions available and welcome all experience levels! 

We offer generous relocation benefits for eligible candidates.

In support of work/life balance, many positions are available for a flexible schedule within the pay period. Ask us about the opportunity for flex scheduling if that’s of interest to you.

Responsibilities:

Lead and mentor an interdisciplinary team consisting of both developers and researchers. The team's core focus is the implementation of ETL pipelines to support a variety of AI/ML and LLM solutions, which in turn address a broad range of customer challenges.

  • Assembles large, complex sets of data to support AI/ML algorithm implementation 
  • ​Builds required infrastructure for optimal extraction, transformation and loading of data from various data sources 
  • ​​Curate and maintain data that is stored in support of metrics and evaluation 
  • ​Implement Artificial Intelligence/Machine Learning algorithms 
  • ​Identifies, designs, and implements internal process improvements including re-designing infrastructure for greater scalability, optimizing data delivery, and automating manual processes 
  • Using Agile methodologies to develop software.​

Required Qualifications:

  • B.S. in data science, AI/ML, computer science, or related field​
  • ​​Minimum six (6) years of relevant experience as a Data Engineer/Scientist.​
  • ​​Experience developing data pipelines and normalizing data with canonical Python packages (e.g. NumPy, Pandas, Polars) 
  • ​Experience contributing on a team using version control (e.g. git, GitLab, Bitbucket)
  • Active TS/SCI U.S. Government Security Clearance with a recent Full-Scope Polygraph (FSP)​

Preferred Qualifications:

  • M.S. or PhD in data science, AI/ML, computer science, or related field​
  • ​Experience with Gitlab, DevSecOps utilizing test-driven development, containers, (e.g. Docker, Docker Compose), cloud services (e.g. AWS), tools for distributed computing (e.g. Spark, Pyspark)  
  • Experience leading an interdisciplinary team of researchers and software developers
  • Experience with any of the following:
    • Large Language Models and experience identifying ways to incorporate them into new domains and applications
    • Applying Transformer-based architectures to domains in other areas outside of Natural Language Processing (NLP) such as computer vision
    • Natural Language Processing algorithms such as BERT
    • Reinforcement learning and familiarity with Gymnasium Gym, OpenEnv, TorchRL, RLlib, and Stable Baselines
    • Applying clustering algorithms and/or deep neural networks to real life problems
    • Implementing tracking and pattern-of-life algorithms
  • Experience with GenAI Ops techniques (e.g. LLM-as-a-judge) and frameworks (e.g. LangFuse, MLFlow, Arize Phoenix)
  • Experience with Machine Learning libraries and frameworks such as HuggingFace and LangChain
  • Experience with Linux
  • Familiarity with using AWS cloud computing resources such as EC2, S3, Lambda, Bedrock, etc.
  • Experience with any of the following additional languages: Java, C++, Rust, Go, and/or C#
  • Experience implementing algorithms on the GPU in Python or C++ using CUDA and other CUDA libraries
  • Experience in application deployment, virtualization, and containerization (e.g. Podman, Docker, Kubernetes, Rancher)
  • Experience shaping and writing proposals

Location: Herndon, VA

Herndon offers a charming blend of small-town ambiance and modern conveniences. In historic Herndon you'll find small town charm while maintaining all the amenities of big city life. The area has beloved farms to explore, historic sites, festivals, local parks, breweries, and some of the best restaurants in the area. Its proximity to D

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

ARKA Group, LP

View company profile →