Best Data Engineering Tools for GitHub

Find and compare the best Data Engineering tools for GitHub in 2026

Use the comparison tool below to compare the top Data Engineering tools for GitHub on the market. You can filter results by user reviews, pricing, features, platform, region, support options, integrations, and more.

  • 1
    AnalyticsCreator Reviews
    See Tool
    Learn More
    AnalyticsCreator serves as a design application centered around metadata, specifically tailored for teams working in data engineering within the Microsoft ecosystem. Engineers can establish structures, transformation processes, loading logic, and dependencies in a centralized manner, allowing for the automatic generation of native SQL, SSIS, Azure Data Factory, Microsoft Fabric, and Power BI components. This approach fosters repeatable methods for data ingestion, transformation, historical data management, slowly changing dimensions (SCD) processing, and deployment, significantly minimizing manual engineering efforts while ensuring that lineage, documentation, and change impact are seamlessly integrated with the overall project design.
  • 2
    Domo Reviews
    Top Pick
    Domo has become a part of Progress Software, integrating its AI and data platform into Progress' suite of offerings. This cloud-native AI data readiness platform not only enhances but also expands Progress' existing data solutions, fostering significant synergies that facilitate the development of innovative, secure, and scalable AI data readiness solutions on a global scale. These combined strengths will assist clients in transforming fragmented enterprise data and insights into governed, AI-ready intelligence, thus elevating the security, governance, and cost-effectiveness of AI-driven projects. Positioned as the agentic platform for the intelligent enterprise, Domo empowers organizations to connect, govern, activate, and disseminate both data and AI effectively. Collaborating with cloud data platforms such as Snowflake, BigQuery, and Databricks, Domo enables the conversion of governed data into various AI agents, applications, workflows, dashboards, and analytics. Additionally, its robust data foundation, activation, and distribution layers empower teams to create and implement intelligence precisely where their work occurs, all while ensuring governance, security, and access controls are firmly in place across both data and AI initiatives. This comprehensive approach not only streamlines operations but also enhances the overall effectiveness of data-driven decision-making within organizations.
  • 3
    Fivetran Reviews
    Fivetran is a comprehensive data integration solution designed to centralize and streamline data movement for organizations of all sizes. With more than 700 pre-built connectors, it effortlessly transfers data from SaaS apps, databases, ERPs, and files into data warehouses and lakes, enabling real-time analytics and AI-driven insights. The platform’s scalable pipelines automatically adapt to growing data volumes and business complexity. Leading companies such as Dropbox, JetBlue, Pfizer, and National Australia Bank rely on Fivetran to reduce data ingestion time from weeks to minutes and improve operational efficiency. Fivetran offers strong security compliance with certifications including SOC 1 & 2, GDPR, HIPAA, ISO 27001, PCI DSS, and HITRUST. Users can programmatically create and manage pipelines through its REST API for seamless extensibility. The platform supports governance features like role-based access controls and integrates with transformation tools like dbt Labs. Fivetran helps organizations innovate by providing reliable, secure, and automated data pipelines tailored to their evolving needs.
  • 4
    Iterative Reviews
    AI teams encounter obstacles that necessitate the development of innovative technologies, which we specialize in creating. Traditional data warehouses and lakes struggle to accommodate unstructured data types such as text, images, and videos. Our approach integrates AI with software development, specifically designed for data scientists, machine learning engineers, and data engineers alike. Instead of reinventing existing solutions, we provide a swift and cost-effective route to bring your projects into production. Your data remains securely stored under your control, and model training occurs on your own infrastructure. By addressing the limitations of current data handling methods, we ensure that AI teams can effectively meet their challenges. Our Studio functions as an extension of platforms like GitHub, GitLab, or BitBucket, allowing seamless integration. You can choose to sign up for our online SaaS version or reach out for an on-premise installation tailored to your needs. This flexibility allows organizations of all sizes to adopt our solutions effectively.
  • 5
    Mozart Data Reviews
    Mozart Data is the all-in-one modern data platform for consolidating, organizing, and analyzing your data. Set up a modern data stack in an hour, without any engineering. Start getting more out of your data and making data-driven decisions today.
  • 6
    Chalk Reviews

    Chalk

    Chalk

    Free
    Experience robust data engineering processes free from the challenges of infrastructure management. By utilizing straightforward, modular Python, you can define intricate streaming, scheduling, and data backfill pipelines with ease. Transition from traditional ETL methods and access your data instantly, regardless of its complexity. Seamlessly blend deep learning and large language models with structured business datasets to enhance decision-making. Improve forecasting accuracy using up-to-date information, eliminate the costs associated with vendor data pre-fetching, and conduct timely queries for online predictions. Test your ideas in Jupyter notebooks before moving them to a live environment. Avoid discrepancies between training and serving data while developing new workflows in mere milliseconds. Monitor all of your data operations in real-time to effortlessly track usage and maintain data integrity. Have full visibility into everything you've processed and the ability to replay data as needed. Easily integrate with existing tools and deploy on your infrastructure, while setting and enforcing withdrawal limits with tailored hold periods. With such capabilities, you can not only enhance productivity but also ensure streamlined operations across your data ecosystem.
  • Previous
  • You're on page 1
  • Next