Principal Data Engineer
AvidXchange, Inc. · Charlotte, North Carolina, United States
- Senior
- Full-time
- Posted 2026-09-23
- Confirmed live on 25 September 2026
Job description
Overview
The Principal Data Engineer is a senior technical leader within AvidXchange's Data Engineering organization responsible for architecting, building, and scaling modern data platforms. In this role, you will drive the migration of legacy Azure SQL Server workloads to Databricks, design real-time streaming pipelines with Apache Kafka, enable AI and agentic capabilities such as Databricks Genie, and define the long-term data architecture strategy. You will partner closely with Software Engineering, Product, Architecture, DevOps, and Analytics teams to deliver secure, high-performing, and reliable data solutions at enterprise scale.
What You'll Do
Data Platform Architecture & Modernization
· Lead the design and implementation of scalable, cloud-native data architectures on Databricks (Delta Lake, Unity Catalog, Lakehouse patterns).
· Own and execute the migration strategy from legacy Azure SQL Server to Databricks, including schema translation, ETL/ELT re-platforming, data validation, and cutover planning.
· Define data modeling standards (medallion architecture, star/snowflake schemas) and ensure consistency across all pipelines and domains.
· Evaluate and recommend tools, frameworks, and platforms to support long-term data strategy and organizational goals.
· Collaborate with Solution and Enterprise Architects to review and approve new data architecture designs.
Streaming & Real-Time Data Engineering
· Architect and implement Kafka-based streaming pipelines for real-time data ingestion, transformation, and delivery.
· Design event-driven architectures and streaming topologies using Kafka Streams, ksqlDB, or Spark Structured Streaming on Databricks.
· Establish patterns for schema management (Confluent Schema Registry), consumer group strategy, offset management, and dead-letter queuing.
· Ensure streaming pipelines meet SLA requirements for latency, throughput, and fault tolerance.
Optimization, Quality & Standards
· Debug and optimize Spark jobs, Delta Lake tables, and SQL workloads for performance, cost efficiency, and maintainability.
· Lead code reviews focused on senior engineers to enforce standards, best practices, and technical quality.
· Manage pipeline quality, data models, and CI/CD delivery workflows; guide teams on continuous improvement.
· Promote strong data management practices — data quality, lineage, observability, and governance.
· Identify opportunities to improve service delivery methods, processes, and resource utilization.
Leadership, Mentorship & Strategy
· Mentor data engineers at all levels, with particular emphasis on developing senior talent.
· Establish and evolve data engineering standards, best practices, and management of technical debt.
· Stay current on data platform trends (Databricks releases, Kafka ecosystem, open table formats) and contribute to long-term architectural vision.
· Develop plans for data security, disaster recovery, backup, business continuity, and archiving across the Lakehouse.
AI, Agentic Capabilities & Intelligent Data Products
· Design and enable AI and agentic capabilities on the Databricks platform, including Databricks Genie for natural language data exploration and self-service analytics.
· Architect data foundations — clean, governed, well-documented Delta tables — that power Genie spaces, AI/BI dashboards, and LLM-driven data agents.
· Collaborate with ML and AI teams to build and maintain feature stores, vector stores, and retrieval-augmented generation (RAG) pipelines on Databricks.
· Evaluate and integrate emerging agentic frameworks (LangChain, Mosaic AI Agent Framework) to automate data workflows and enable intelligent data products.
· Define governance and observability standards for AI-driven data pipelines, ensuring reliability, auditability, and responsible AI practices.
Cross-Functional Collaboration
· Partner with project managers and business leaders on initiatives involving enterprise data.
· Collaborate across teams to influence and strengthen data engineering practices organization-wide.
· Work with Analytics, ML, and product engineers to design and deliver end-to-end data solutions that meet business needs.
What We're Looking For
Required
· Bachelor's degree in Computer Science, Engineering, or a related field with 10+ years of data engineering experience in a high-availability, business-critical environment.
· Hands-on Databricks expertise: Delta Lake, Unity Catalog, Databricks Workflows, Databricks SQL, and cluster/job optimization.
· Proven experience migrating from legacy Azure SQL Server (or other relational RDBMS) to a Databricks Lakehouse — schema translation, data validation, and cutover strategies.
· Strong proficiency in Apache Kafka for streaming pipelines — producers/consumers, topic design, partitioning strategy, and Kafka Connect.
· Expert-level PySpark and/or Scala Spark skills, including performance tuning, broadcasting, partitioning, and caching.
· Deep understan
Prepare for the interview
Nothing collected for this employer yet. The Blind 75 is what technical screens draw from; practise it here, with a coach, in Java or Python.
More at AvidXchange, Inc.
- Senior Marketing Campaign Manager · Charlotte, North Carolina, United States
- Product Marketing Manager II, GTM Operations · Virtual
- Vice President of Financial Planning & Analysis · Charlotte, North Carolina, United States
- Brand Marketing Manager II · Charlotte, North Carolina, United States
- Senior Director of Business Development · Charlotte, North Carolina, United States
- Account Executive · Virtual
- Strategic Partnerships Business Development Representative II · Charlotte, North Carolina, United States; Virtual
- Sr. Revenue Enablement Manager, Channel Sales · Virtual