Senior Software Engineer (Data Platform)
Databricks · Bengaluru, India · Engineering - Pipeline · listed September 23, 2024
The shape of it
Seniority
Senior
Experience asked
4–6 years
Where
Not stated
Requirements listed
7
Length
836 words
In the posting’s own words
As a Senior Software Engineer working on the Data Platform team you will help build the Data Intelligence Platform for Databricks that will allow us to automate decision-making across the entire company. You will achieve this in collaboration with Databricks Product Teams, Data Science, Applied AI and many more. You will develop a variety of tools spanning logging, orchestration, data transformation, metric store, governance platforms, data consumption layers etc. You will do this using the latest, bleeding-edge Databricks product and other tools in the data ecosystem - the team also functions as a large, production, in-house customer that dog foods Databricks and guides the future direction of the product.
What it asks for · 7
- 6+ years of industry experience
- 4+ years of experience providing technical leadership on large projects similar to the ones described above - ETL frameworks, metrics stores, infrastructure management, data security.
- Experience building, shipping and operating reliable multi-geo data pipelines at scale.
- Experience working with and operating workflow or orchestration frameworks, including open source tools like Airflow and DBT or commercial enterprise tools.
- Experience with large-scale messaging systems like Kafka or RabbitMQ or commercial systems.
- Excellent cross-functional and communication skills, consensus builder.
- Passion for data infrastructure and for enabling others by making their data easier to access.
What the job covers
- Design and run the Databricks metrics store that enables all business units and engineering teams to bring their detailed metrics into a common platform for sharing and aggregation, with high quality, introspection ability and query performance.
- Design and run the cross-company Data Intelligence Platform, which contains every business and product metric used to run Databricks. You’ll play a key role in developing the right balance of data protections and ease of shareability for the Data Intelligence Platform as we transition to a public company.
- Develop tooling and infrastructure to efficiently manage and run Databricks on Databricks at scale, across multiple clouds, geographies and deployment types. This includes CI/CD processes, test frameworks for pipelines and data quality, and infrastructure-as-code tooling.
- Design the base ETL framework used by all pipelines developed at the company.
- Partner with our engineering teams to provide leadership in developing the long-term vision and requirements for the Databricks product.
- Build reliable data pipelines and solve data problems using Databricks, our partner’s products and other OSS tools. Provide early feedback on the design and operations of these products.
- Establish conventions and create new APIs for telemetry, debug, feature and audit event log data, and evolve them as the product and underlying services change.
- Represent Databricks at academic and industrial conferences & events.
Tools and skills named
Data
- Databricks12×
- Data pipelines2×
- ETL2×
- Airflow
Security & compliance
- Security3×
- Audit
Cloud & infra
- CI/CD
- Kafka
- RabbitMQ
Ways of working
- Cross-functional
Words the posting leans on
- data21×
- platform9×
- product9×
- experience5×
- metrics5×
- scale5×
- tools5×
- customers4×
- design4×
- frameworks4×
- infrastructure4×
- operating4×
- pipelines4×
- run4×
- build3×
- data intelligence3×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.