PyJobs PyJobs
Post a job
Company
Magpie Health Analytics logo

Magpie Health Analytics

magpie-health-analytics.breezy.hr
Location
Fully remote
Apply

Data Engineer

Position Purpose:

We are seeking a Data Engineer to support the development of cloud-native data solutions that improve operational efficiency, support regulatory reporting, and drive actionable insight for our commercial and government healthcare clients. This role is responsible for building, maintaining, and optimizing scalable data pipelines and validation workflows across complex datasets—enabling downstream analytics, application logic, and system modernization initiatives.

Primary Duties & Responsibilities:

  • Design, develop, and maintain ETL/ELT pipelines using structured and semi-structured data from relational databases, flat files, APIs, and cloud data sources.

  • Collaborate with backend and architecture teams to define data transformation flows aligned with dashboard and reporting application needs.

  • Design and optimize data schemas to ensure performance, integrity, and compatibility with reporting requirements.

  • Develop and maintain efficient, testable, and reusable data processing scripts using Python, SQL, and cloud-native tools.

  • Collaborate with DevOps, analysts, and application developers to align pipelines with system architecture, storage strategy, and reporting needs.

  • Implement data quality and validation checks and document data lineage and pipeline logic for audit and reuse.

  • Troubleshoot performance issues in data jobs and support data pipeline operations across environments (DEV, VAL, PROD).

  • Contribute CI/CD workflows and automation strategies to promote rapid iteration and secure deployment of data services.

  • Assist in developing or maintaining data documentation, including metadata, data dictionaries, and technical user guides.

  • Stay up to date with emerging technologies, techniques, and trends to inform product development and decision-making.

Minimum Qualifications:

  • Bachelor's degree in computer science, engineering, statistics, or related field.

  • 4+ years of experience in a data engineering, data pipeline, or ETL/ELT development role.

  • Strong proficiency in SQL, data transformation logic, and performance tuning for large datasets.

  • Proficiency with Python and libraries such as Pandas, PySpark, or Numpy.

  • Experience with modern version control and CI/CD practices (e.g., GitHub Actions, Jenkins).

  • Understanding of distributed computing solutions for data processing (e.g. AWS Glue, AWS EMR, Apache Hadoop, Apache Spark)

  • Experience with data pipeline orchestration frameworks or serverless tools (e.g., AWS StepFunctions, AWS Glue Workflows, and/or AWS Lambda).

Preferred Qualifications:

  • Experience developing solutions in AWS cloud environments (S3, Lambda, Aurora, SNS, CloudFormation/CDK).

  • Experience with Snowflake, Amazon RDS, or other cloud-native data warehouses.

  • Familiarity with modern data transformation tools (dbt, Dataform, or equivalent) and best practices for modular, tested SQL transformations within an ELT architecture.

  • Experience supporting healthcare data systems or CMS data environments

  • AWS certifications are a plus.

Get the latest Python jobs in your inbox.
Email address
Frequency
Receive jobs daily
Best when actively looking for a job
Receive jobs weekly
Best when just browsing jobs

Show filters

Your email won't be used for commercial purposes. Read our Privacy Policy.