Latka logo

Top 8 Synthetic Data Software Companies With $5M–$10M Revenue (August 2026)

As of August 2026, Latka tracks 8 synthetic data software companies with $5M–$10M in annual revenue. They have combined revenues of $47.2M and employ 408 people. They have raised $26.7M.

Every company below sells synthetic data software to other businesses and is ranked by its most recent annual revenue. Revenue, funding, headcount and customer figures come from CEO interviews on the Latka podcast, public company filings, and Latka estimates where a company has not disclosed a number.

What Synthetic Data Software Companies do

Synthetic data software provides tools that generate artificial datasets which mimic real-world data. These datasets can be used for development, testing, and training machine learning models while ensuring privacy and compliance with data protection regulations. The primary use cases include software testing, model training, and data analysis where maintaining confidentiality is paramount. These tools typically offer features such as data generation, data masking, and customization options to create datasets that resemble original data patterns. Common buyer personas for synthetic data software include software developers, data scientists, compliance officers, and IT managers who require secure, scalable solutions to maintain data integrity without sacrificing privacy.

Companies
8
Revenue
$47.2M
Funding
$26.7M
Employees
408

Filters

Sorting: Highest -> Lowest

Filters

Top Synthetic Data Software Companies With $5M–$10M Revenue

Showing 8 of 8 companies ranked by annual revenue.

1Sapien logo
Sapien

San Francisco, California, United States

Sapien is a decentralized data foundry, turning collective human knowledge into enterprise-grade AI training data.

Revenue
$8.6M
Year founded
2023
Team size
58
2zypl.ai logo
zypl.ai

Dubai, United Arab Emirates

Zypl.ai is a technology company that develops artificial intelligence–based solutions for the financial sector. The company focuses on optimizing credit scoring for financial institutions using synthetic data and offers advanced technologies for data analysis and process automation.

Revenue
$6.1M
Year founded
2021
Team size
51
3Duality Technologies logo
Duality Technologies

Hoboken, New Jersey, United States

Duality's breakthrough innovative technologies eliminate the conflict between data protection and business growth and innovation. The Duality Data Analytics and AI platform is built upon advanced encryption methods, hardware technologies, and machine learning techniques that protect sensitive data while in use. Duality is the only multi-PET platform with the ability to combine various technologies to meet the unique needs of sensitive data operations. These guardrails streamline and enhance the data operations critical for data-driven insights and innovations by eliminating bulky, expensive, and limiting processes like data anonymization and tokenization. Traditional data protection methods prevent organizations from truly adopting and leveraging advanced models to their benefit, resulting in restrictive policies like "no sensitive data can be used for model training." With Duality, organizations can confidently customize 3rd party models on their own data without fear of data leaks. Model providers can scale model customization knowing that their proprietary model is never exposed to the customer, preventing competitive intelligence leaks. Financial institutions can turn their manual KYC requests into self-service operations, greatly enhancing the speed and success of these expensive requirements. The benefits of data protection guardrails span from efficiency gains, to unlocking previously inaccessible data, to slashing costs on high-security infrastructure.

Revenue
$5.6M
Year founded
2016
Team size
51
4Carpe Data logo
Carpe Data

Santa Barbara, California, United States

Developer of predictive scoring and data products for insurers designed to provide a holistic view into each risk. The company's products leverages the social web, online content, wearables, connected devices and other forms of next-generation data to assess risk at critical steps in the insurance policy lifecycle, aggregate and assess the social web as well as consolidate and functionalize the next generation of data, enabling insurers to more accurately predict risk and innovate with new products to meet changing customer habits.

Revenue
$5.6M
Year founded
2016
Funding
$26.7M
Team size
108
5Flower logo
Flower

United States

Train AI on distributed data

Revenue
$5.6M
Year founded
2023
Team size
37
6SafeGraph logo
SafeGraph

Denver, Colorado, United States

SafeGraph is a data company. That's it - that's all we do. We predict the past. SafeGraph's mission is to democratize access to data. SafeGraph's five year goal is to be THE source for accurate data about every physical place in the world. SafeGraph builds truth sets for machine learning, deep learning, and AI. SafeGraph is unlocking the world's most powerful data so that machines and humans can answer society's toughest questions.

Revenue
$5.4M
Year founded
2016
Team size
49
7Parallel Domain logo
Parallel Domain

San Francisco, California, United States

Simulation platform for testing, evaluating, analyzing, and training AI models at scale. Ensuring public safety while accelerating autonomy development. #syntheticdata #autonomy #AI #computervision #AV #ADAS #machinelearning #syntheticdatarealimpact

Revenue
$5.4M
Year founded
2017
Team size
49
8SBX Robotics logo
SBX Robotics

Toronto, Canada

Synthetic data for better vision.

Revenue
$5M
Year founded
2020
Team size
5

Frequently asked questions about Synthetic Data Software Companies With $5M–$10M Revenue

How many synthetic data software companies with $5M–$10M in annual revenue are there?

Latka tracks 8 synthetic data software companies with $5M–$10M in annual revenue. Together they generate $47.2M in annual revenue and employ 408 people.

Which synthetic data software company with $5M–$10M in annual revenue is the largest?

Sapien is the largest, with $8.6M in annual revenue, founded in 2023.

How much revenue does a typical synthetic data software company with $5M–$10M in annual revenue make?

The average synthetic data software company in this list makes $5.9M a year, across 8 companies with reported revenue.

Who are the leading Synthetic Data software vendors with $5M–$10M in annual revenue?

Ranked by annual revenue, the leaders are Sapien, zypl.ai, Duality Technologies, Carpe Data and Flower.

How much funding have synthetic data software companies with $5M–$10M in annual revenue raised?

The 8 synthetic data software companies with $5M–$10M in annual revenue tracked here have raised $26.7M in disclosed funding between them.

Related Artificial Intelligence Software categories

Inclusion Criteria

- The software must generate synthetic datasets that replicate the structure and characteristics of real data. - It should provide capabilities for data masking and privacy preservation. - The platform should support various data types including structured data, images, and text. - Tools must allow customization to suit different development and testing requirements. - Solutions should integrate easily with existing development and data science workflows. - Not just a data augmentation tool; it must also create entirely synthetic datasets suitable for training and testing.