TixelJobs
A
Archer56via Greenhouse

Staff Data Engineer

San Jose, California, United StatesPosted 2d ago
Data EngineerStaff+Full-time

Not sure if you're a good fit?

Upload your resume and TixelJobs AI will compare it against Staff Data Engineer at Archer56. Get a match score, missing keywords, and improvement tips before you apply.

Free preview · Your resume stays private

About the Role

Archer is an aerospace company based in San Jose, California building an all-electric vertical takeoff and landing aircraft with a mission to advance the benefits of sustainable air mobility. We are designing, manufacturing, and operating an all-electric aircraft that can carry four passengers while producing minimal noise.

Our sights are set high and our problems are hard, and we believe that diversity in the workplace is what makes us smarter, drives better insights, and will ultimately lift us all to success. We are dedicated to cultivating an equitable and inclusive environment that embraces our differences, and supports and celebrates all of our team members.

Archer is best known for Midnight, our electric aircraft. This role isn't on that program. The AI Products Org is building software for the general aviation industry. As a Staff Data Engineer on the AI Platform team, you will design, build, and operate the data infrastructure that powers large-scale model training and inference. You will own the pipelines, storage systems, and data quality mechanisms that sit upstream of our ML platform, ensuring models train on clean, high-throughput, well-governed data.


What You’ll Do:

  • Design and maintain high-throughput, fault-tolerant ingestion and transformation pipelines that feed training workloads at scale, with a focus on latency, throughput, and correctness.
  • Build and operate the data lakehouse — defining table formats (Iceberg, Paimon, Parquet), partitioning strategies, and compaction policies optimized for ML consumption patterns.
  • Instrument pipelines with data quality checks, lineage tracking, and anomaly detection so that model failures trace back to data problems quickly.
  •  Partner with ML engineers to define feature stores, dataset versioning, and experiment-to-production data contracts; integrate with tools like MLflow for dataset and artifact tracking.
  • Work closely with AI researchers, platform engineers, and software engineers to understand data access patterns, optimize query performance, and unblock training runs.
 
What You Need:
  • 5+ years of professional data engineering experience excluding internships. 
  • BS/MS/PhD in Computer Science, Data Engineering, Software Engineering, or a related field.
  • Hands-on experience building production pipelines with tools like Apache Spark, Flink, Airflow, dbt, or similar batch/streaming frameworks.
  • Deep proficiency with columnar formats (i.e. Parquet), open table formats (Apache Iceberg or Apache Paimon), and object storage systems (S3 or equivalent).
  • Experience with high-throughput messaging systems such as Apache Pulsar or Kafka for real-time data ingestion.
  • Strong SQL skills; experience with StarRocks for large-scale analytical queries and real-time analytics over the lakehouse.
  • Familiarity with AWS data services (S3, Glue, EMR) and containerized workloads (Docker/Kubernetes) in production environments. Airflow/Prefect/Dagster


Bonus Qualifications:
  • Experience with CDC (Change Data Capture) replication from transactional systems (i.e. PostgreSQL → Lakehouse via Debezium or Airbyte).
  • Exposure to audio or time-series data pipelines, including preprocessing for ASR or speech model training.
  • Prior experience in aerospace, aviation, or other safety-critical domains where data lineage and auditability are non-negotiable.

 

About The Team:
Share