Data catalog tools for Amazon Redshift

List of data catalogs tools

Data catalog is a structured collection of data used by an organization. It is a kind of data library where data is indexed, well-organized, and securely stored. Most data catalog tools contain information about the source, data usage, relationships between entities as well as data lineage. This provides a description of the origin of the data and tracks changes in the data to its final form.

Dataedo

Dataedo is an on-premises data catalog & metadata management tool. It allows you to catalog, document, and understand your data with a data dictionary, business glossary, and ERDs. It reads your schema and lets you easily describe each data element with descriptions, business-friendly aliases, and custom fields. It features a data community module, which allows you to crowdsource knowledge about data from everyone in your organization.

Automated Cataloging: Yes
Business Glossary: Yes
Commenting/Community: Yes
Commercial: Commercial
Data Classification: Yes
Data Lineage: Yes
Data Profiling: Yes
Export: HTML,MS Excel,PDF
Free edition: No
Rating of assets: Yes
Dataedo Data Catalog
Dataedo Sensitive Data Discovery
Dataedo Report Catalog
Dataedo Data Search
Dataedo Referenece Data Management
Dataedo Data Catalog - list of data sources
Dataedo Data Community
Dataedo Business Glossary
Dataedo Data Profiling
Dataedo Data Lineage
Dataedo ERD

Alation Data Catalog

Alation pioneered the data catalog market and is now leading its evolution into a platform for a broad range of data intelligence solutions including data search & discovery, data governance, stewardship, analytics, and digital transformation. Thanks to its powerful Behavioral Analysis Engine, inbuilt collaboration capabilities, and open interfaces, Alation combines machine learning with human insight to successfully tackle even the most demanding challenges in data and metadata management.

More than 250 enterprises realize business outcomes with Alation, including Salesforce, Cisco, Docusign, Finnair, Pfizer, Nasdaq, and Albertsons.

Automated Cataloging: Yes
Business Glossary: Yes
Commenting/Community: Yes
Commercial: Commercial
Data Classification: No
Data Lineage: Yes
Data Profiling: Yes
Export: MS Excel
Free edition: No
Rating of assets: Yes

Informatica Enterprise Data Catalog

Informatica Data Catalog is a machine learning-based data catalog that lets you classify and organize data assets across any environment to maximize data value and reuse, and provides a metadata system of record for the enterprise. It automatically scans and catalogs data across the enterprise, indexing it for enterprise-wide discovery using simple, Google-like search.

Automated Cataloging: Yes
Business Glossary: Yes
Commenting/Community: Yes
Commercial: Commercial
Data Classification: Yes
Data Lineage: Yes
Data Profiling: Yes
Export: CSV,JSON,MS Excel,Plain text,XML
Free edition: No
Rating of assets: Yes

Lumada Data Catalog

Lumada Data Catalog software leverages AI, machine learning, and patented fingerprinting technology to automate the discovery, classification, and management of your enterprise data. It simplifies access and promotes collaboration allowing an organization to more intelligently use their data.

Automated Cataloging: Yes
Business Glossary: Yes
Commenting/Community: Yes
Commercial: Commercial
Data Classification: Yes
Data Lineage: Yes
Data Profiling: Yes
Export: XML
Free edition: No
Rating of assets: Yes

OvalEdge

OvalEdge is a data catalog tool that automatically organizes and catalogs your data using machine learning and advance algorithms. You can organize data using tags, usage statistics, user names, and other markers – so it’s easily retrievable with everyday language.

Automated Cataloging: Yes
Business Glossary: Yes
Commenting/Community: Yes
Commercial: Commercial
Data Classification: Yes
Data Lineage: Yes
Data Profiling: Yes
Export: MS Excel
Free edition: No
Rating of assets: No

Alteryx Connect

Alteryx Connect is a social data cataloging and data exploration platform for the enterprise. The powerful data cataloging provided by Alteryx Connect centralizes business terms and definitions, metrics, and information assets for maximum consistency, discoverability, and collaboration.

Automated Cataloging: Yes
Business Glossary: Yes
Commenting/Community: Yes
Commercial: Commercial
Data Classification: Yes
Data Lineage: Yes
Data Profiling: Yes
Export: CSV,MS Excel,PDF
Free edition: No
Rating of assets: Yes

Truedat

Truedat is an open source data cataloging and governance tool that allows to quickly unify and explore combined metadata from different sources on the same interface. It enables to organize & enrich information through configurable workflows and monitor data governance activity.

Automated Cataloging: Yes
Business Glossary: Yes
Commenting/Community: Yes
Commercial: Free
Data Classification: Yes
Data Lineage: Yes
Data Profiling: Yes
Export: CSV
Free edition: Yes
Rating of assets: No

Global IDs

The Global IDs Data Catalog automates the linking of logical business data models to physical data assets, keeps the metadata up to date, and scales with the size of your enterprise, from small to very large.

Automated Cataloging: Yes
Business Glossary: Yes
Commenting/Community: -
Commercial: Commercial
Data Classification: Yes
Data Lineage: Yes
Data Profiling: Yes
Export: XML
Free edition: No
Rating of assets: No

Tree Schema

The Tree Schema data catalog provides all of the essential catalog capabilities including rich-text documentation, data lineage, assigning data stewards and technical owners to your data assets, tagging your assets and much more. You can point Tree Schema to your database and fully populate your catalog in under 5 minutes. Tree Schema also supports non-traditional data sources including S3, Kafka and DynamoDB.

Automated Cataloging: Yes
Business Glossary: Yes
Commenting/Community: Yes
Commercial: Commercial
Data Classification: Yes
Data Lineage: Yes
Data Profiling: Yes
Export: Plain text
Free edition: Yes
Rating of assets: No
Data Asset Expert
Comments
Data schema overview

Atlan

Atlan is a modern, cloud native data catalog. It's ease of use and intuitive interface enables diverse personas including engineers, data stewards and business users to discover, understand and trust data. Atlan leverages machine learning and a bots ecosystem to automate documentation and stewardship tasks such as automatic data profiling, data quality alerts and glossary tagging. It is built on an Open API architecture, and has a pay as you go pricing model, making it a good fit for teams of all sizes.

Automated Cataloging: Yes
Business Glossary: Yes
Commenting/Community: Yes
Commercial: Commercial
Data Classification: Yes
Data Lineage: Yes
Data Profiling: Yes
Export: -
Free edition: No
Rating of assets: Yes

Stemma

Stemma is a fully managed data catalog, powered by the leading open-source data catalog, Amundsen. By bridging the gap between data producers and data consumers, Stemma enables you to gain total trust in your data. Stemma provides enterprise management (easy deployment, enterprise-grade security) and richer metadata. It makes finding trustworthy data easy and offers an always up-to-date view of your data's usage at any time through automated documentation based on common usage patterns.

Automated Cataloging: Yes
Business Glossary: Yes
Commenting/Community: Yes
Commercial: Commercial
Data Classification: No
Data Lineage: Yes
Data Profiling: Yes
Export: -
Free edition: No
Rating of assets: Yes

Select Star

Select Star automatically catalogs & documents your database tables and BI dashboards. You can find out where your data is coming from, which dashboards are built on top of it, who is using the data, and how they are using it.

Automated Cataloging: Yes
Business Glossary: Yes
Commenting/Community: Yes
Commercial: -
Data Classification: Yes
Data Lineage: Yes
Data Profiling: No
Export: CSV,JSON
Free edition: Yes
Rating of assets: No

Data catalogs are part of data management tools. They enable automatic metadata management with user-friendly form that makes data easy to understand even for non-IT members of the organisation.

The key feature of data catalogs is to provide metadata context to the user in a way that allows different teams within the organization (both IT and Non-IT) to discover and understand relevant data.

From the organization's perspective, the important functions of data catalog tools are also:
• storage of data resources from different repositories as well as from different engine systems - compatibility with multiple connectors,
• automation of data management processes,
• advanced resource search by name, type, date of change, owner, etc.
• data lineage,
• automated data Classification,
• Discovering data relationship and dependencies between objects,
• Business Glossary, unifying nomenclature and definitions of terms,
• Data Profiling,

Data stewards, business teams, and data analysts often struggle with the problem of what specific data means, where it comes from, and which elements it is directly related to. These are just a few problems for which Data catalog tools have been created. Based on the imported repositories, data catalogs enable automated cataloging and organizing of data, solving the problem of time-consuming querying of the resources.

To avoid misunderstandings data catalog tools provide a Business Glossary, through which the nomenclature is systematized. It contains business terms along with their definition, relationship to each other, as well as its location in the hierarchy of all data assets.

There are many apps for data catalog tasks on the market. We have listed complex data cataloging software that can also solve data profiling, data lineage, and data classification problems, as well as open-source data catalog tools.