Latka logo

Top 2 Machine Learning Data Catalog Software Companies With $1M–$5M Revenue (September 2026)

As of September 2026, Latka tracks 2 machine learning data catalog software companies with $1M–$5M in annual revenue. They have combined revenues of $3.3M and employ 52 people.

Every company below sells machine learning data catalog software to other businesses and is ranked by its most recent annual revenue. Revenue, funding, headcount and customer figures come from CEO interviews on the Latka podcast, public company filings, and Latka estimates where a company has not disclosed a number.

What Machine Learning Data Catalog Software Companies do

Machine Learning Data Catalog Software serves as a specialized framework designed to enhance the management, discovery, and utilization of data specifically for machine learning projects. These solutions facilitate real-time data discovery by automating the cataloging of datasets, enabling organizations to effectively organize and manage their data assets. In doing so, they allow data scientists and machine learning engineers to locate relevant datasets quickly, thereby accelerating the development of machine learning models. Typical features of Machine Learning Data Catalog Software include automated metadata ingestion, lineage tracking, and advanced search capabilities powered by machine learning algorithms. This facilitates easier dataset evaluation and improves collaboration across teams, as stakeholders can access effective data documentation and understand the provenance of their data. Common buyer personas include data scientists, machine learning engineers, data governance professionals, and IT managers, all of whom seek efficient ways to manage and utilize large volumes of data for analytical and operational purposes.

Companies
2
Revenue
$3.3M
Funding
-
Employees
52

Filters

Sorting: Highest -> Lowest

Filters

Top Machine Learning Data Catalog Software Companies With $1M–$5M Revenue

Showing 2 of 2 companies ranked by annual revenue.

1Gensyn logo
Gensyn

London, United Kingdom

Gensyn is a machine learning computing protocol that connects all of the machine learning-capable compute hardware to enable the training of deep learning models. It focuses on decentralizing computing power to advance the field of machine learning.

Revenue
$2.2M
Year founded
2020
Team size
44
2DreamForTek logo
DreamForTek

Leiria, Portugal

DreamForTek is a technology and engineering company specializing in turnkey automation, robotics, and industrial digital solutions. The firm focuses on helping manufacturing and industrial clients modernize and optimize their production processes through advanced automation technologies, robotics integration, artificial vision systems, and Industry 4.0 / IoT solutions.

Revenue
$1.1M
Year founded
2019
Team size
8

Frequently asked questions about Machine Learning Data Catalog Software Companies With $1M–$5M Revenue

How many machine learning data catalog software companies with $1M–$5M in annual revenue are there?

Latka tracks 2 machine learning data catalog software companies with $1M–$5M in annual revenue. Together they generate $3.3M in annual revenue and employ 52 people.

Which machine learning data catalog software company with $1M–$5M in annual revenue is the largest?

Gensyn is the largest, with $2.2M in annual revenue, founded in 2020.

How much revenue does a typical machine learning data catalog software company with $1M–$5M in annual revenue make?

The average machine learning data catalog software company in this list makes $1.7M a year, across 2 companies with reported revenue.

Who are the leading Machine Learning Data Catalog software vendors with $1M–$5M in annual revenue?

Ranked by annual revenue, the leaders are Gensyn and DreamForTek.

Related IT Infrastructure Software categories

Inclusion Criteria

- Must offer automated metadata management to simplify data organization - Should provide advanced search functionalities to enable quick data discovery - Must include lineage tracking to visualize data flow and relationships - Should facilitate collaboration among teams by offering clear documentation and accessibility - Must cater specifically to machine learning use cases, not just general data management - Should integrate seamlessly with existing data tools and platforms used by the organization - Not just a data storage solution; must actively support data discovery and utilization features