Jobs and Careers
EX

Principal AI Engineer

Exactera
UKRemotefull_timeVerifiedPosted 19 May 2026

About the role

Exactera has offices in New York City, Tarrytown NY, San Diego, CA, London, and Argentina. 

The Role

As Principal AI Engineer, you will own the inference and intelligence layer of the Exactera platform. You will build the substrate that agentic workflows run on: the domain knowledge graph and structured representations of expert reasoning, the hybrid retrieval that operates over them, model serving and LLM integration, the agent platform, the evaluation harness that captures expert judgment as ground truth, and the interfaces that product engineers compose into customer-facing workflows.

You report directly to the CTO and work closely with the data engineering team (close collaboration on data structure, entity resolution, and the contract between the lakehouse and the intelligence layer), product engineering (who consume your interfaces to build agentic workflows), product management (who set product direction and prioritize the workflows the platform supports), and domain experts in tax advisory (who provide the ground truth and judgment the AI systems are evaluated against — and whose decisions you will help capture as first-class data).

This is an individual contributor role with architectural authority over the AI/ML platform. You will set technical direction, make build-versus-buy decisions, and provide direction to other senior engineers working on platform components. You will have meaningful input on stack evolution and are expected to evaluate the stack against real workloads and propose changes when warranted.

 

What You Will Build

We have made initial choices you will inherit and refine. The data platform runs on Databricks with Unity Catalog for governance. We use MLflow for experiment tracking and model lifecycle. Our LLM integrations use Anthropic and OpenAI APIs. MCP is our current pattern for exposing capabilities to agentic workflows, with production gateways already in service. We are on AWS, with Terraform for infrastructure-as-code. Exact tool experience matters less than having strong, defensible opinions about the categories.

 

Knowledge and Reasoning Substrate

This is the headline of the role. Tax-advisory work is relationship-heavy: companies own subsidiaries, subsidiaries transact, transactions map to jurisdictions, jurisdictions have regulatory frameworks, and comparable selections are justified against multi-dimensional functional profiles. A vector store can suggest candidates; it cannot defend a selection under audit. The substrate you build is what makes the rest of the platform compliance-grade.

  • Domain ontology and knowledge graph. Design and operate the typed graph of tax-domain entities and relationships — companies, transactions, jurisdictions, segments, functional profiles, comparability factors, expert decisions, and reports — with relationships reified (e.g., comparable-to as an edge carrying which dimensions, who weighted them, in which report, with what outcome). Stack choice is yours.
  • Entity resolution and relationship extraction at scale. Pipelines that resolve entities across 34,000 reports, SEC/EDGAR filings, and customer data into a canonical, versioned representation. Relationship extraction from free text into structured edges. Reconciliation against curated reference data (NAICS, ownership filings).
  • Expert judgment as first-class data. Capture practitioner decisions — selections, rejections, weightings, and the reasoning behind them — as structured entities attached to the graph, versioned and queryable. This is Exactera's proprietary moat encoded.
  • Versioned graph snapshots. The graph evolves; audit defense requires being able to reconstruct exactly what it looked like at any point in time. Design the versioning and snapshot strategy.

 

Hybrid Retrieval

Retrieval is three modes — graph traversal, vector search, and structured queries — and a planner that decides which combination to use per task. "Find comparable companies for this intercompany loan" is a different retrieval shape than "summarize prior treatment of intercompany IP licensing in EMEA." The retrieval layer makes the right combination of moves automatically.

  • The full RAG pipeline, as one retrieval mode within the larger system: chunking strategies, embedding generation, index management, retrieval optimization, and context assembly for LLM consumption. Embedding pipelines for heterogeneous data, with index maintenance as source data and the graph evolve.
  • Graph-enhanced retrieval. Graph traversal for relationship-aware lookups, graph-guided chunking that respects entity boundaries, graph context assembly that pulls in related entities and prior precedents alongside narrative text.<

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Exactera

View company profile →