Jobs and Careers
ON
Software Engineer
OnehouseUnited Statesfull_timeVerifiedPosted 22 Feb 2024
About the role
About OnehouseOnehouse is a mission-driven company dedicated to freeing data from data platform lock-in. We deliver the industry’s most interoperable data lakehouse through a cloud-native managed service built on Apache Hudi. Onehouse enables organizations to ingest data at scale with minute-level freshness, centrally store it, and make available to any downstream query engine and use case (from traditional analytics to real-time AI / ML).
We are a team of self-driven, inspired, and seasoned builders that have created large-scale data systems and globally distributed platforms that sit at the heart of some of the largest enterprises out there including Uber, Snowflake, AWS, Linkedin, Confluent and many more. Riding off $33M total funding and a fresh Series A backed by Greylock/Addition, we are quickly expanding and looking for rising talent to grow with us and become future leaders of the team. Come help us build the world's best fully managed and self-optimizing data lake platform!
The Community You Will JoinWhen you join Onehouse, you're joining a team of passionate professionals tackling the deeply technical challenges of building a 2-sided engineering product. Our engineering team serves as the bridge between the worlds of open source and enterprise: contributing directly to and growing Apache Hudi (already used at scale by global enterprises like Uber, Amazon, ByteDance etc) and concurrently defining a new industry category - the transactional data lake. The Data Infrastructure team is the grounding heartbeat to all of this. We live and breathe databases, building cornerstone infrastructure by working under Hudi's hood to solving incredibly complex optimization and systems problems.
We are a team of self-driven, inspired, and seasoned builders that have created large-scale data systems and globally distributed platforms that sit at the heart of some of the largest enterprises out there including Uber, Snowflake, AWS, Linkedin, Confluent and many more. Riding off $33M total funding and a fresh Series A backed by Greylock/Addition, we are quickly expanding and looking for rising talent to grow with us and become future leaders of the team. Come help us build the world's best fully managed and self-optimizing data lake platform!
The Community You Will JoinWhen you join Onehouse, you're joining a team of passionate professionals tackling the deeply technical challenges of building a 2-sided engineering product. Our engineering team serves as the bridge between the worlds of open source and enterprise: contributing directly to and growing Apache Hudi (already used at scale by global enterprises like Uber, Amazon, ByteDance etc) and concurrently defining a new industry category - the transactional data lake. The Data Infrastructure team is the grounding heartbeat to all of this. We live and breathe databases, building cornerstone infrastructure by working under Hudi's hood to solving incredibly complex optimization and systems problems.
The Impact You Will Drive:
- As a software engineer at Onehouse, you will contribute directly to Apache Hudi and the surrounding open source ecosystem, while deploying and operating these technologies at massive scale for our customers.
- Accelerate our open source <> enterprise flywheel by working on the guts of Apache Hudi's transactional engine and optimizing it for diverse Onehouse customer workloads.
- Act as a SME to deepen our teams' expertise on database internals, query engines, storage and/or stream processing.
A Typical Day:
- Build systems that enable users to manage petabytes of data with a fully managed cloud service.
- Build functionality that enables data systems to be cloud native (self managed), scalable (auto scaling) and secure (different levels of access control).
- Build scalable job management on Kubernetes to ingest, store, manage and optimize petabytes of data on cloud storage.
- Design systems that help scale and streamline metadata and data access from different query/compute engines.
- Exhibit full ownership of product features, including design and implementation, from concept to completion.
- Be passionate about designing for future scale and high availability, while possessing a deep understanding of common failure patterns and their remediations.
- Uphold a high engineering bar around the code, monitoring, operations, automated testing, release management of the platform.
What You Bring to the Table:
- 4+ years of experience as a software engineer with experience developing distributed systems.
- Strong, object-oriented design and coding skills with Java.
- Experience with inner workings of distributed (multi-tiered) systems, algorithms, and relational databases.
- Deal well with ambiguous/undefined problems; ability to think abstractly; articulate technical challenges and solutions.
- Speed and hustle → Ability to prioritize across feature development and tech debt.
- Ability to solve complex programming/optimization problems.
- Ability to quickly prototype optimization solutions and analyze large/complex data.
- Clear communication skills.
- Experience working on database systems, Query Engines or Spark codebases.
- Experience working on cloud based (data focused) services.
- Deep understanding of Spark, Flink, Presto, Hive, Parquet internals.
- Hands-on experience with open source projects like Hadoop, Hive, Delta Lake, Hudi, Nifi, Drill, Pulsar, Druid, Pinot, etc.
Nice to haves (but not required):
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s