Staff Data Engineer (Agent Systems)
Faraday FutureAbout the role
The Company:
AIxCrypto is a U.S.-listed company dedicated to building a world-leading ecosystem that integrates AI and blockchain while bridging Web2 and Web3. Its core products include the BesTrade DeAI Agent and the AIxC ecosystem products. Our mission is to build the core infrastructure of an AI-driven global economy that empowers more intelligent, more transparent, and more efficient capital appreciation and value regeneration. Our vision is to create a world where real-world value flows freely, securely, and intelligently across information networks — empowering global users to effortlessly participate in the co-creation and sharing of value within the crypto economy.
Your Role:
As a Staff Data Engineer (Agent Systems) in our Crypto projects, you will design, deliver, and operate the data platform that powers our agentic products—real-time ingestion (on-chain/market/social), feature stores, vectorization & retrieval for RAG, time-series/streaming computation, and ML observability. You’ll define data contracts and SLAs, ensure offline–online consistency, and partner closely with AI Agents, Backend/BFF, and Security/Compliance.
Key Responsibilities:
- Platform Architecture: Author data contracts and schemas; produce ADRs; design tiered storage (OLTP/OLAP, lakehouse) with governance and lineage.
- Streaming & Batch Pipelines: Build low-latency streams (Kafka/Flink or equivalent) and robust batch ETL (Airflow/Dagster + dbt); support CDC, replay/backfill, and schema evolution.
- Feature Store & Online Serving: Provide point-in-time-correct features (near-real-time/time-series); guarantee offline–online parity and latency SLOs.
- RAG Data Plane: Orchestrate embedding pipelines, chunking/routing, vector DB (pgvector/FAISS/Milvus), HNSW/IVF indexes, and reindexing/TTL strategies.
- Evaluation & ML Ops: Materialize canonical eval datasets/labels; wire A/B hooks; manage model/feature registries and CI/CD for ML; enable canary rollouts.
- Data Quality & Observability: Monitor freshness, completeness, duplication, drift/decay; implement lineage and cost/performance guardrails.
- Security & Compliance: Enforce PII handling, retention, and auditability; implement least-privilege access to datasets and secrets.
- Collaboration: Work hand-in-hand with DS/Agents/Backend on interfaces, SLAs, and incident RCAs; document playbooks and standards.
Basic Qualifications:
- Bachelor’s degree or above in CS/EE/Math/Stats or related.
- 7+ years in data engineering with 3+ years building streaming pipelines/feature stores for production systems.
- Proficient in Python and SQL, plus one of Java/Scala/Go, strong data modeling and performance tuning.
- Streaming & batch: Kafka, Flink/Spark (stateful ops, event-time, watermarking, exactly-once), Airflow/Dagster, dbt.
- Storage: PostgreSQL/MySQL, ClickHouse/BigQuery/Snowflake, NoSQL (MongoDB/DynamoDB/Bigtable/Firestore), Redis, and lakehouse on Amazon S3 or Google Cloud Storage (GCS) (Parquet + Iceberg/Delta).
- Feature platforms (e.g., Feast) and online feature serving; offline–online consistency validation.
- Vector retrieval: embeddings pipelines and vector stores (pgvector/FAISS/Milvus); relevance & recency metrics.
- Ops/Observability: Docker/K8s; data quality/lineage (OpenLineage/Marquez or similar); cost & throughput optimization.
Salary Range:
($160,000 - $200,000 DOE), plus benefits and incentive plans
Preferred Qualifications:
- Crypto-signal ingestion (order books, trades, on-chain events); precision arithmetic and idempotent metrics.
- Privacy/compliance (GDPR/CCPA), tokenization/pseudonymization strategies.
- Cost/perf tuning (autoscaling, comp
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s