Senior Databricks Data Engineer– COBOL/Mainframe Modernization
*Clearance*
Active Public Trust clearance is required.
*Company Overview*
Spearmint Solutions is a technology-focused consulting and engineering company delivering scalable, high-performance software solutions for commercial clients and government organizations. We specialize in modern application architectures, cloud and data modernization, automation, and enterprise software development.
Our teams work collaboratively to design, build, modernize, and support mission-critical applications while emphasizing clean engineering practices, reliability, performance, and practical solutions to complex business challenges.
*About The Role*
*We are seeking an experienced Senior Databricks Data Engineer with COBOL and mainframe modernization experience* to support the migration and transformation of legacy enterprise data workloads into a modern cloud-based lakehouse architecture.
This role combines modern data engineering with legacy-system expertise. The engineer will work with *Databricks, Apache Spark, Python, SQL, Delta Lake, cloud storage, and enterprise ETL pipelines* while analyzing legacy COBOL programs, copybooks, record layouts, batch-processing logic, and mainframe data structures.
The ideal candidate can bridge the gap between legacy business logic and modern data platforms—understanding how data is represented and processed in COBOL/mainframe environments and translating those workloads into scalable, maintainable Databricks pipelines.
*Responsibilities*
* Design and develop scalable *ETL/ELT pipelines using Databricks and Apache Spark*.
* Modernize legacy mainframe data-processing workloads into cloud-based Databricks solutions.
* Analyze *COBOL applications, copybooks, file definitions, record layouts, and batch-processing logic*.
* Translate legacy business rules and data transformations into modern Python, PySpark, SQL, and Databricks workflows.
* Parse and transform fixed-width, delimited, hierarchical, and other legacy data formats.
* Support conversion of mainframe datasets into modern relational, lakehouse, and analytical data structures.
* Develop reliable ingestion pipelines for high-volume enterprise data.
* Build and maintain *Delta Lake* tables and medallion-style lakehouse architectures.
* Develop data-validation and reconciliation processes to ensure migrated data accurately reflects source-system information.
* Implement data-quality controls, logging, exception handling, and operational monitoring.
* Optimize Spark jobs, Databricks clusters, queries, storage layouts, and pipeline performance.
* Develop reusable Python, PySpark, and SQL components for enterprise data processing.
* Support batch and incremental data-processing workflows.
* Collaborate with legacy-system SMEs, data architects, application developers, testers, and business stakeholders.
* Troubleshoot complex data-transformation and migration defects.
* Participate in source-to-target mapping, technical design, and data-modeling activities.
* Support CI/CD and version-controlled deployment of Databricks notebooks, jobs, and data pipelines.
* Maintain technical documentation covering source layouts, transformation rules, data lineage, and pipeline architecture.
*Required Qualifications*
* *7+ years* of professional data engineering, ETL, database development, or related software engineering experience.
* *3+ years* of hands-on experience developing production workloads using *Databricks*.
* Strong experience with *Apache Spark / PySpark*.
* Strong *Python and SQL* development skills.
* Experience designing and maintaining large-scale *ETL/ELT data pipelines*.
* Strong knowledge of *Delta Lake and lakehouse architecture concepts*.
* Experience with data modeling, transformation, validation, reconciliation, and data-quality engineering.
* Experience processing high-volume enterprise datasets.
* Hands-on experience working with *COBOL applications, COBOL copybooks, or mainframe data structures*.
* Ability to interpret legacy file definitions, field layouts, data types, and business-processing logic.
* Experience translating legacy data-processing rules into modern data pipelines.
* Understanding of batch-processing architectures and enterprise scheduling/workflow concepts.
* Experience with Git and modern CI/CD development practices.
* Strong troubleshooting and performance-optimization skills.
* Strong communication skills and ability to collaborate with both legacy-system and modern cloud engineering teams.
*Preferred Qualifications*
* Active Public Trust and existing GFE access.
* *Databricks Certified Data Engineer Associate or Professional* certification.
* Experience with *AWS-based Databricks environments*.
* Experience with AWS S3 or comparable cloud object storage.
* Experience with *Unity Catalog* and Databricks governance capabilities.
* Experience with Databricks Workflows, Auto Loader, Delta Live Tables/Lakeflow, or Structured Streaming.
* Experience working with *EBCDIC, VSAM, DB2, flat-file, fixed-width, or mainframe extract formats*.
* Experience reverse-engineering COBOL batch applications or legacy transformation logic.
* Experience with large-scale *mainframe-to-cloud or legacy-to-modern data modernization programs*.
* Experience supporting federal or other highly regulated enterprise environments.
* Familiarity with PostgreSQL, NoSQL databases, Kafka, or downstream analytics platforms.
*Pay Range: $110,000 – $150,000 USD*
Pay: $110,000.00 - $150,000.00 per year
License/Certification:
* active Public Trust Clearance (MBI) (Required)
Work Location: Remote