Data profiling tools for MySQL
Data Profiling tools allow analyzing, monitoring, and reviewing data from existing databases in order to provide critical insights. Data profiling can help organizations improve data quality and decision-making process by identifying problems and addressing them before they arise.
Kylo
Kylo is an open source enterprise-ready data lake management software platform. It lets you search and explore data and metadata, view lineage, and profile statistics. In addition, it offers self-service data ingest with data cleansing, validation, and automatic profiling.
Access control: | |
---|---|
Commercial: | Free |
Desktop/Cloud: | Cloud |
Excel workbooks: | |
Flat files: | |
Free edition: | |
Metadata identification: | |
NoSQL sources: | |
Runs on: (for desktop): | - |
Sensitive data discovery: | |
SQL sources: | |
Statistics of data: | Avg,Max,Min,Stdev |
Tagging data: |
Astera Centerprise
Astera Centerprise is an end-to-end data integration software that enables you to integrate, cleanse, and transform data in a code-free environment. Its built-in data profiling feature lets you easily examine your source data and get detailed information about its structure, quality, and integrity. Custom data integration and quality rules can also be defined to validate incoming data and identify missing or invalid records.
Access control: | |
---|---|
Commercial: | Commercial |
Desktop/Cloud: | Desktop |
Excel workbooks: | |
Flat files: | |
Free edition: | |
Metadata identification: | |
NoSQL sources: | |
Runs on: (for desktop): | Windows |
Sensitive data discovery: | |
SQL sources: | |
Statistics of data: | Avg,Max,Min |
Tagging data: |
DataCleaner
The heart of DataCleaner is a strong data profiling engine for discovering and analyzing the quality of your data. Find the patterns, missing values, character sets and other characteristics of your data values.
Access control: | |
---|---|
Commercial: | Free |
Desktop/Cloud: | Desktop |
Excel workbooks: | |
Flat files: | |
Free edition: | |
Metadata identification: | |
NoSQL sources: | |
Runs on: (for desktop): | Linux,Mac OS,Windows |
Sensitive data discovery: | |
SQL sources: | |
Statistics of data: | - |
Tagging data: |
CloverDX
CloverDX Data Profiler is a CloverDX module that lets you perform various analyses of your data. It is a part of CloverDX Designer and helps to do various profiling tasks, such as finding the maximum value, median, the most unique value, and many others.
Access control: | |
---|---|
Commercial: | Commercial |
Desktop/Cloud: | Desktop |
Excel workbooks: | |
Flat files: | |
Free edition: | |
Metadata identification: | |
NoSQL sources: | |
Runs on: (for desktop): | Linux,Mac OS,Windows |
Sensitive data discovery: | |
SQL sources: | |
Statistics of data: | Avg,Max,Min,Stdev |
Tagging data: |
Trifacta
Trifacta is an open and interactive cloud platform for data engineers and analysts to collaboratively profile, prepare, and pipeline data for analytics and machine learning. For ease of data profiling, Trifacta automatically identifies dataset formats, schemas, specific attributes, and relationships across attributes and datasets, along with associated metadata for each dataset.
Access control: | |
---|---|
Commercial: | Commercial |
Desktop/Cloud: | Cloud |
Excel workbooks: | |
Flat files: | |
Free edition: | |
Metadata identification: | |
NoSQL sources: | |
Runs on: (for desktop): | - |
Sensitive data discovery: | |
SQL sources: | |
Statistics of data: | Avg,Max,Min,Stdev |
Tagging data: |
DQLabs
DQLabs platform has a data profiling platform that is AI-driven and accepts data from multiple sources in different formats if necessary. The user interface is user-friendly and will allow the user to track the data profiling process and make adjustments where they feel it’s necessary. The platform algorithms will detect deep insight into the source data and increase the quality of the profiled data.
Access control: | |
---|---|
Commercial: | Commercial |
Desktop/Cloud: | Cloud |
Excel workbooks: | |
Flat files: | |
Free edition: | |
Metadata identification: | |
NoSQL sources: | |
Runs on: (for desktop): | - |
Sensitive data discovery: | |
SQL sources: | |
Statistics of data: | - |
Tagging data: |
SAS Data Quality
SAS Data Quality gives you a single interface to manage the entire data quality life cycle: profiling, standardizing, matching, and monitoring. It lets you validate data against standard measures and customized business rules. Uncover relationships across tables, databases, and source applications. Verify that the data in your tables matches the appropriate description. Establish trends and commonalities in business information and examine numerical trends via mean, median, mode, and standard deviation.
It makes it easy to profile and identify problems, preview data, and set up repeatable processes to maintain a high level of data quality.
Access control: | |
---|---|
Commercial: | Commercial |
Desktop/Cloud: | Cloud |
Excel workbooks: | |
Flat files: | |
Free edition: | |
Metadata identification: | |
NoSQL sources: | |
Runs on: (for desktop): | - |
Sensitive data discovery: | |
SQL sources: | |
Statistics of data: | Avg,Stdev |
Tagging data: |
StarDQ
StarDQ is a powerful enterprise solution for profiling, cleansing, augmenting, and standardizing the data to significantly improve returns on corporate intelligence initiatives.
Access control: | |
---|---|
Commercial: | Commercial |
Desktop/Cloud: | Cloud |
Excel workbooks: | |
Flat files: | |
Free edition: | |
Metadata identification: | |
NoSQL sources: | |
Runs on: (for desktop): | - |
Sensitive data discovery: | |
SQL sources: | |
Statistics of data: | - |
Tagging data: |
Alation Data Catalog
Alation’s data profiling capabilities help reduce the time spent in the data exploration phase. With Alation’s data profile, data consumers have the metrics they need to easily discern the quality of any data object. Alation displays important characteristics, statistics, and numerical graphs about the data — enabling data scientists and data engineers to quickly take action. The data profiling now also includes new charts and customizations.
Access control: | |
---|---|
Commercial: | Commercial |
Desktop/Cloud: | Cloud |
Excel workbooks: | |
Flat files: | |
Free edition: | |
Metadata identification: | |
NoSQL sources: | |
Runs on: (for desktop): | - |
Sensitive data discovery: | |
SQL sources: | |
Statistics of data: | - |
Tagging data: |
The use of data profiling tools can lead to higher-quality, more reliable data or eliminating errors that add costs to data-driven projects. Eliminating these costly errors involve processes such as:
• Collecting descriptive statistics.
• Collecting data types, length and recurring patterns.
• Tagging data with keywords, descriptions or categories.
• Performing data quality assessment.
• Discovering metadata and assessing its accuracy.
The most efficient way of handling the data profiling process is to automate it with a data management solution. We prepared a list of open-source data profiling tools that help you carry out the analysis of your data and identify the issues.