Data Engineering Startups funded by Y Combinator (YC) 2026 | Y Combinator
Data Engineering Startups funded by Y Combinator (YC) 2026
September 2026
Browse 87 of the top Data Engineering startups funded by Y Combinator.
We also have a Startup Directory where you can search through over 5,000 companies.
Fivetran\ \ Y Combinator LogoW2013\ \ • Active • 1,200 employees • San Francisco
Fivetran automates data movement out of, into and across cloud data platforms. We automate the most time-consuming parts of the ELT process from extracts to schema drift handling to transformations, so data engineers can focus on higher-impact projects with total pipeline peace of mind. With 99.9% uptime and self-healing pipelines, Fivetran enables hundreds of leading brands across the globe, including Autodesk, Conagra Brands, JetBlue, Lionsgate, Morgan Stanley, and Ziff Davis, to accelerate data-driven decisions and drive business growth. Fivetran is headquartered in Oakland, California, with offices around the world.
data-engineering
saas
analytics
b2b
CueBench\ \ Y Combinator LogoS2026\ \ • Active • 3 employees • San Francisco
data-engineering
reinforcement-learning
Magma\ \ Y Combinator LogoS2026\ \ • Active • 1 employees • San Francisco
We enable companies monetize their agents' traces.
ai
reinforcement-learning
data-engineering
Hebbian Robotics\ \ Y Combinator LogoS2026\ \ • Active • 2 employees • San Francisco
Hebbian is an open source SDK for building quality control pipelines for Physical AI. We enable data teams to scale their pipeline from thousands to millions of hours of data, without managing infrastructure.
https://github.com/Hebbian-Robotics/hflow
robotics
data-engineering
data-science
databases
artificial-intelligence
Olam Labs\ \ Y Combinator LogoS2026\ \ • Active • 2 employees • San Francisco
Multi-agent environments let us evaluate and train models in complex, simulated worlds. We work with researchers on both evaluations and training for character and agentic performance typically difficult to assess for using current popular datasets.
Play one of our first releases, Multi-Agent Arena (https://olamlabs.ai/arena), where humans come play social strategy games against multiple AI agents. It's currently top 50 on OpenRouter and has thousands of matches played each day. We use Multi-Agent Arena to build datasets on agentic performance and real-world socialization.
reinforcement-learning
gaming
data-engineering
artificial-intelligence
Alkera AI\ \ Y Combinator LogoS2026\ \ • Active • 3 employees • San Francisco
Alkera is a data engineering/analysis/science agent that works in your IDE or CLI. If you’re tired of agents that produce false, or even dangerous, reports about your data, then Alkera is for you. Our agent is equipped with your data’s history and lineage, your team’s knowledge, and tailored plugins for data apps such as Snowflake and dbt. Simply install and connect your services to begin doing trusted agentic data work. We also offer VPC & on-prem deployment options for when data can't leave your boundary.
data-science
data-engineering
developer-tools
enterprise-software
artificial-intelligence
Litmus\ \ Y Combinator LogoS2026\ \ • Active • 3 employees • New York City
Litmus is building the most accurate framework for evaluating and benchmarking human capability, starting with software.
AI will compound small differences in human capability into increasingly large differences in what people can accomplish, while making existing static benchmarks obsolete. Software is already there: AI can hill-climb any output-based evaluation, while the ability to direct it is becoming the defining advantage.
Litmus applies the same approach we already use for model evals to humans – creating world-like environments, and inspecting trajectory instead of just output.
Every knowledge industry will soon face the same problem. We build Litmus to tell you what humans are capable of.
We're already helping build frontier technical teams at Mercor, Composio, Neo Scholars, and more.
recruiting
data-engineering
ai
Petrarch\ \ Y Combinator LogoS2026\ \ • Active • 3 employees • San Francisco
Petrarch brings internal company data to frontier labs, starting with manufacturing and industrial companies. We source data including codebases, project files, and payments from bankruptcy courts, then de-identify and prepare the data for model training.
b2b
marketplace
data-engineering
Markov\ \ Y Combinator LogoS2026\ \ • Active • 2 employees • San Francisco
Markov sells expert computer-use data to frontier labs. We’ve sold more than 33k+ hours and our open-source datasets have got over 200k+ downloads on hugging face.
reinforcement-learning
data-labeling
data-engineering
Corvera\ \ Y Combinator LogoW2026\ \ • Active • 4 employees • San Francisco
Corvera (YC W26) is the context layer for AI-native CPG brands. We help brands unlock the full potential of AI by making their unified data legible to any AI tool via MCP.
As of mid-March 2026, we scaled from $0 to $33k in MRR in 4 weeks, are serving 12 brands, and are growing 130% week-on-week.
With Corvera in place, brands can enable anyone in their organization to: - Rapidly deploy AI agents via tools such as Claude and ChatGPT, - Build dashboards and apps using Cursor and Lovable, and - Automate workflows across the entire business, from category management to supply chain.
All in a secure, audit-logged, and access-controlled environment, without the need for a team of data engineers or AI specialists, in a simple plug-and-play solution.
We're founded by Chris (2x Founder; ex-CEO at Better Nature; Forbes 30U30), Dirk (ex-Google Data & AI Lead), and Matthew (ex-Head of Product at Rosemark; Princeton MEng in CS).
To-date, we’ve raised $6.2M and are backed by YC, firstminute capital, 6 Degrees Capital, 20VC, Rebel Fund, Duke Capital Partners, and more. Our ICP are omnichannel retail-focused CPG brands at the $100M+ revenue inflection point.
Our mission is bold: to enable the next generation of unicorn CPG brands to be run by teams of less than 10 people.
Our vision is even bolder: to be the system of intelligence for the $5.4T global CPG industry.
Visit www.corvera.ai to learn more!
b2b
saas
data-engineering
consumer
ai
Velum Labs\ \ Y Combinator LogoW2026\ \ • Active • 2 employees • San Francisco
Velum is the operating system for data quality. Velum automatically monitors and enforces data quality across a company's data stack, so bad data never reaches dashboards. We turn data quality from a manual task into infrastructure that runs itself.
Data trust you can prove. From the pipeline to the boardroom.
machine-learning
data-engineering
Captain\ \ Y Combinator LogoW2026\ \ • Active • 2 employees • San Francisco
Captain syncs complex, multimodal files and cloud storage from sources like S3, SharePoint, and Google Drive, and makes that knowledge searchable for your agents.
Instead of stitching together parsers, embeddings, vector databases, rerankers, and retrieval infrastructure, Captain self-tunes the retrieval pipeline to your data and workloads.
On benchmarks, this improves accuracy from roughly 78% with standard RAG to over 97% by tuning retrieval to the unique structure and content of your data.
Captain is purpose-built for complex, regulated, and high-stakes file search workloads.
Learn more at https://captain.dev
data-engineering
infrastructure
b2b
api
Velvet\ \ Y Combinator LogoF2025\ \ • Active • 4 employees • San Francisco
Infra and data for interactive AI.
ai
data-engineering
conversational-ai
generative-ai
Hyperspell\ \ Y Combinator LogoF2025\ \ • Active • 8 employees • San Francisco
Hyperspell is your Company Brain. AI agents are brilliant and clueless. They ace any test and still have no idea how your company works. Hyperspell connects your tools and synthesizes documents and conversations into a live, permissioned context graph. Any agent can read from it and write back to it like a filesystem, with every fact traced to source.
ai
machine-learning
saas
data-engineering
DeepAware (Robotics Center of Silicon Valley, https://roboticscenter.ai/) is the fastest way for enterprises and researchers to get robots and robotics parts in the US — 72-hour delivery or Bay Area pickup.
Beyond hardware, we help teams collect teleoperation data, build reinforcement learning environments, and deploy robots into production. Customers include AI labs, industrial operations, research teams, and event producers.
supply-chain
data-engineering
robotics
machine-learning
artificial-intelligence
sieve\ \ Y Combinator LogoP2025\ \ • Active • 2 employees • New York City
sieve solves data cleaning for hedge funds and investment firms by letting them get clean data in four lines of code. Currently, their data pipelines have conditions that raise for human review, which literally send an email to engineers with data that needs to be reviewed. We provide an API that integrates directly into their existing pipeline - instead of raising for human review, they can send all the same information to our API and get clean, high-quality data back.
By using our AI agents built specifically for financial data collection, along with expert-in-the-loop review, we provide our clients with clean, validated data at a scale and level of quality that wasn't achievable before.
apis
investing
data-engineering
Vision Lab\ \ Y Combinator LogoP2025\ \ • Active • 11 employees • San Francisco
We capture and structure real factory workflows at scale by combining first-person industrial video with SOP-level process knowledge.
This enables robotics and AI labs to train on real production data, not just controlled lab environments.
manufacturing
ai
robotics
data-engineering
Nitrode\ \ Y Combinator LogoW2025\ \ • Active • 10 employees • San Francisco
Nitrode is a research company focused on improving game development with AI. To advance game development, we believe AI models firstly need to be judged against robust evaluation frameworks that reflect the actual complexity of the field.
artificial-intelligence
data-engineering
machine-learning
b2b
Kilvin\ \ Y Combinator LogoF2024\ \ • Active • 3 employees • New York City
Kilvin is an AI-native agency that builds custom software solutions for non-technical teams.
With AI, the future of software is democratized. Every team deserves tools built exactly for them, not off-the-shelf SaaS products they need to adapt to. Kilvin makes it a reality, delivering a private, secure workspace of custom applications, built around your workflows and ready in days, not months.
No engineers needed on your side. Just software that fits exactly how you work.
developer-tools
automation
ai
data-engineering
productivity
Melder\ \ Y Combinator LogoF2024\ \ • Active • 2 employees • New York City
Melder is an Excel add-in that brings AI functions and document support into your spreadsheets. Upload files directly into cells, use smart formulas like =GEN, and build automations—all without leaving Excel.
Core features: - File-to-Sheet: Drop PDFs directly into cells, then reference them in formulas. - AI-Powered Functions: Write formulas like =GEN() or =EXTRACT() to summarize, classify, and analyze content. - Chat Assistant: Use our AI assistant to help build your sheets or answer questions from your data, live in the workbook.
Business users use Melder to: - Accelerate diligence by extracting insights from data rooms - Review contracts by identifying key terms and clauses instantly - Run market research by pulling information from competitor websites - Synthesize transcripts by generating summaries from interviews and calls
Melder brings the power of structured spreadsheet logic to the messy, unstructured data world—no coding needed.
data-engineering
generative-ai
artificial-intelligence
Sensei\ \ Y Combinator LogoS2024\ \ • Active • 2 employees • Boston
Sensei helps robotics companies scale and outsource their training data collection. Our hardware platform enables the collection of human-demonstration data at a tenth of the cost and twice the speed of current teleop approaches. Our software platform acts like Scale AI for robotics data: a large network of paid human operators use our low-cost collection platform to fulfill data-generation requests.
robotics
hard-tech
artificial-intelligence
marketplace
data-engineering
Overstand Labs\ \ Y Combinator LogoW2025\ \ • Active • 4 employees • New York City
Overstand is a data lab that allows our customers to navigate any set of data in just a few minutes.
*Enterprise*: For enterprises, we we unify Slack, email, calls, and operational data, then surface the signals that matter — customer needs, risks, and revenue opportunities hidden in everyday conversations.
*Legal Firms*: For legal firms, we help them really quickly understand their entire discovery corpus (either before, or after document review), and quickly build out an initial case assessment and facts.
Instead of waiting on reports or manual analysis, teams get immediate, evidence-backed answers from the data they already have. Overstand delivers clarity and leverage as your business scales.
data-science
data-engineering
conversational-ai
artificial-intelligence
legaltech
Sharpe\ \ Y Combinator LogoS2024\ \ • Active • 3 employees • San Francisco
Sharpe helps traders go from idea to profit in minutes with AI, bundling petabytes of market data with high-performance infrastructure.
ai
finance
analytics
data-engineering
big-data
Zeit AI\ \ Y Combinator LogoS2024\ \ • Active • 12 employees • Munich, Germany
Ask a question in plain English and ZeitMind connects SAP, your ERP, CRM, HR systems and 600+ other sources, prepares the messy data, and turns it into the live applications your teams actually make decisions with.
It is built for the finance, controlling, procurement, supply chain and sales teams at companies that run enterprise systems without an enterprise data team.
artificial-intelligence
analytics
b2b
saas
data-engineering
Mica AI\ \ Y Combinator LogoS2024\ \ • Active • 3 employees • San Francisco
Mica's AI agents replace the data ops teams fixing bad data.
When bad or missing data breaks the pipeline, and orchestration, retries, and monitoring fail, painful manual review work kicks in, pulling humans in to investigate and patch data issues across systems. Mica does what those humans do: gathering the right information from internal docs and external systems, reasoning across context, and resolving errors autonomously to get the pipeline moving again.
The result: dramatically reduce time, cost, and operational drag as your data pipelines scale without scaling ops headcount. Mica turns judgment-heavy data fixes from a manual bottleneck into an automated background process with full auditability.
data-engineering
data-science
enterprise-software
Trellis AI\ \ Y Combinator LogoW2024\ \ • Active • 34 employees • San Francisco
b2b
data-engineering
databases
infrastructure
ai
kater.ai\ \ Y Combinator LogoW2024\ \ • Active • 3 employees • San Francisco
1. You explain your problem. 2. Kater identifies the most important data questions to ask. 3. Kater writes the code. 4. You get insights in seconds rather than weeks.
Kater.ai flips the script on enterprise analytics by making every user an expert analyst. It uses a continuous classification engine to turn a single business question into a contextualized package of questions that is specific to your needs.
Kater puts the power of data into the hands of business experts while ensuring they use trusted data that is specific to their persona. No more waiting for data analysts. No more wasted time on analysis misfires and rework.
Yvonne was a data engineer and analyst who built the entire data stack at CREXi. Robin led engineering in Microsoft.
Data is the new oil. Companies are data-rich, insight-poor. We're helping companies become insight-rich. This is the future of data.
data-engineering
analytics
artificial-intelligence
Ocular AI\ \ Y Combinator LogoW2024\ \ • Active • 8 employees • San Francisco
ai
data-engineering
machine-learning
speech-recognition
Reducto\ \ Y Combinator LogoW2024\ \ • Active • 30 employees • San Francisco
Reducto is the agentic document platform for leading AI teams building systems to automate document workflows at enterprise scale.
Our platform provides a comprehensive toolkit for working with documents the way a human would, combining custom in-house and leading frontier models to power efficient and accurate document workflows. Reducto is trusted by leading AI teams at companies like Harvey, Scale AI, and Vanta.
We are built for enterprise workloads with flexible deployment options from the cloud to fully air-gapped environments, SOC II and HIPAA compliance, and zero data retention.
Learn more: https://reducto.ai/ Find our YC deal: https://reducto.ai/yc
documents
data-engineering
search
enterprise-software
ai
DataShare\ \ Y Combinator LogoS2023\ \ • Active • 1 employees • Austin, TX, USA
DataShare is a data-as-a-service platform that lets you embed charts, dashboards and exports directly into your product. For example, if you run an accounting startup, DataShare would enable you to embed a full profit and loss dashboard, with downloadable statements. DataShare is backed by an enterprise-grade data warehouse, and can be implemented in fewer than 20 lines of code.
data-engineering
databases
analytics
Cedalio\ \ Y Combinator LogoS2023\ \ • Active • 6 employees • San Francisco
Cedalio is the AI agent platform for the Office of the CFO. We automate procurement, AP, and the data flows behind them — from invoice intake to 3-way match to payment — so finance teams stop doing the work and start governing it.
data-engineering
ai
finance
fintech
Artie\ \ Y Combinator LogoS2023\ \ • Active • 16 employees • San Francisco
Artie is software that streams data from databases to data warehouses in real-time. Today, most companies run their ETL process every few hours or overnight, so their data warehouse is always out of date; with Artie, the warehouse always has live production data.
data-engineering
developer-tools
enterprise-software
saas
Ohm\ \ Y Combinator LogoW2023\ \ • Active • 6 employees • San Francisco
Ohm’s AI platform helps Fortune 100 engineering teams accelerate the engineering and testing of complex hardware, including wearables, electric vehicles, and batteries. Launched in 2025, Ohm's platform is used by leading teams to compress development cycles, optimize test programs, and launch better products, faster.
data-engineering
ai
artificial-intelligence
hardware
TableFlow\ \ Y Combinator LogoW2023\ \ • Active • 2 employees • San Francisco
TableFlow builds AI teammates for data tasks, helping operations and data teams automate the messy, manual tasks buried in PDFs, spreadsheets, images, and emails.
ai
automation
documents
data-engineering
saas
Honeydew\ \ Y Combinator LogoW2023\ \ • Active • 6 employees • Tel Aviv-Yafo, Israel
The way people use data is constantly changing. Data teams must support every new context without breaking the shared truth. Honeydew’s semantic layer does it automatically. We validate each change and update every data flow.
Using Honeydew, data teams can support 10x more data users - without more engineers or compromising integrity.
saas
data-engineering
analytics
b2b
Sunpia\ \ Y Combinator LogoS2022\ \ • Active • 3 employees • San Francisco
Sunpia lets developers easily experience the cost and speed benefits of serverless infrastructure, without having to rewrite their code. Developers annotate their code and Sunpia automatically designs a microservice version of it they can deploy on their own cloud.
data-engineering
developer-tools
kubernetes
Midplane\ \ Y Combinator LogoS2022\ \ • Active • 2 employees • Berlin, Germany
Companies are connecting AI coding agents like Cursor and Claude directly to their databases — where one wrong command can delete data or leak customer records. Midplane sits in between and controls what each agent is allowed to do: it blocks dangerous queries, hides sensitive data, and keeps a record of everything the agent touches. It’s open source, and runs either as a hosted service or on your own servers.
b2b
productivity
generative-ai
data-engineering
ai
MovingLake\ \ Y Combinator LogoS2022\ \ • Active • 3 employees • Mexico City, CDMX, Mexico
MovingLake is Fivetran for event-driven architectures. Companies such as Casai use our product to obtain orders and price changes in real time.
analytics
api
data-engineering
b2b
saas
Lamin\ \ Y Combinator LogoS2022\ \ • Active • 10 employees
Improve your workflows with open-source context and data access.
Query, trace, and govern with a lineage-native, format-agnostic lakehouse for agents and teams.
With support for biological formats and registries — by the creators of Scanpy
biotech
data-engineering
machine-learning
developer-tools
open-source
Sanifu\ \ Y Combinator LogoS2022\ \ • Active • 5 employees • Nairobi, Kenya
Sanifu automates the repetitive work spreadsheets are used for. Teams across finance, marketing, operations, and revenue simply describe their repetitive work like cleaning and validating data, matching records, preparing reports and moving data from emails, PDFs and Excel exports into dashboards and other systems. Sanifu captures the steps, rules and outputs so the work is repeated reliably every time
Some of these work never needed spreadsheets to begin with; the likes of data cleanup, bank reconciliations, report consolidation, payment allocations, channel attributions. Spreadsheets have become the standard workbench for handling these tasks and in the process, teams lose valuable time doing manual data entry, filling templates, formatting data, applying formulas, deleting rows or columns.
Sanifu is replacing the spreadsheet while maintaining the logic and processes behind the work, getting you from inputs to your desired outputs in seconds. This significantly reduces time spent on copy-paste work, spreadsheet cleanup, and manual entries and checking, while improving accuracy, consistency, and operational visibility. Sanifu helps teams scale critical back-office processes without relying on brittle spreadsheets, scripts, or general AI tools that reinvent the logic every run.
Sanifu is a full operational platform with 5k+ app integrations, 1k+ native functions, forms, scheduler, data stores, dashboards, approvals, reviews, AI-powered entity matching among many other features that make it an end-to-end data app management platform
b2b
saas
data-engineering
analytics
ai
LanceDB\ \ Y Combinator LogoW2022\ \ • Active • 35 employees • San Francisco
LanceDB is a new open-source vector database that can support low-latency billion-scale vector search on a single node. Built around a new columnar data format, LanceDB makes it incredibly easy to build applications for generative AI, recsys, search engines, content moderation, and more.
aiops
data-engineering
machine-learning
open-source
Elementary\ \ Y Combinator LogoW2022\ \ • Active • 12 employees • Tel Aviv-Yafo, Israel
Elementary enables data teams to detect problems in their data before their users do. An open-source solution that any data engineer can deploy in minutes without sharing sensitive data.
data-engineering
analytics
developer-tools
open-source
Dynamo AI\ \ Y Combinator LogoW2022\ \ • Active • 40 employees • San Francisco
End-to-end privacy, security, and compliance solutions to prepare your organization for emerging AI regulations.
privacy
data-engineering
machine-learning
Sieve\ \ Y Combinator LogoW2022\ \ • Active • 34 employees • San Francisco
Sieve builds the data and environments frontier AI labs use to train the next generation of multimodal systems.
AI is moving beyond chatbots into video, audio, images, software, robotics, and interactive worlds. The next generation of models will need to understand how the world looks, sounds, moves, responds, and changes over time. Progress is bottlenecked by one thing: high-quality data.
Sieve brings together exabyte-scale infrastructure, novel multimodal understanding techniques, large-scale sourcing, and deep research partnerships to create datasets and environments with unmatched precision, quality, and speed. This has earned the trust of frontier AI labs, Fortune 100 companies, and fast-growing AI startups working on generative media, robotics, computer use, world models, and agentic systems.
video
developer-tools
data-engineering
data-labeling
ai
Versable\ \ Y Combinator LogoW2022\ \ • Active • 3 employees • San Francisco
Auto parts retailers get product data from hundreds of manufacturers that is inaccurate and inconsistent, often with big gaps in key values. They currently have a team of "catalog managers" who are required to process and enhance this data line by line, resulting in a week to months long lag between receiving product data and actually being able to start generating revenue from those products.
Versable leverages AI to scan the web for tens of millions of auto parts listings, and uses a fine-tuned LLM with RAG to instantly process, enhance, and transform data. With just a part number, Versable is able to generate market-ready titles, product descriptions, and specs, in any format that's needed.
automotive
manufacturing
data-engineering
ai
Trackingplan\ \ Y Combinator LogoW2022\ \ • Active • 17 employees • Barcelona, Spain
Trackingplan automatically discovers and monitors all the information your applications and websites are collecting, ensuring that you can trust your BI, analytics, marketing, and sales tools.
You can think of us as Segment Protocols but totally transparent, where developers can keep using Google Analytics, Amplitude, Hubspot, Intercom, Braze, etc. as they are used to.
Installed in minutes in using your Tag Manager or adding just one line of code to your web or apps, we model all the data being sent to third parties. Since Trackingplan understands what each piece of data means, it identifies patterns, detects anomalies, and automatically connects the dots to create value from data that was hidden in plain sight:
- An always up-to-date single source of truth and data governance tool. To discover, understand and document your data and improve communication across teams. - Automated notifications when something breaks or changes. To make sure that integrations are always well implemented: Schema errors, traffic anomalies, rogue events... - Easy to understand, customizable, cross-service alerts. To detect trends, insights, and problems without using complex, engineer-oriented solutions.
data-engineering
analytics
saas
Pipekit\ \ Y Combinator LogoS2021\ \ • Active • 9 employees
Our app manages Argo Workflows for data teams, enabling complex data & CI pipelines in half the time while saving companies hundreds of thousands of dollars annually. We maintain Argo Workflows, an open-source pipeline framework for Kubernetes that’s used in production by Bloomberg, Intuit, Adobe, New Relic, NVIDIA, and many other open-source early adopters.
open-source
developer-tools
data-engineering
devops
Evidence\ \ Y Combinator LogoS2021\ \ • Active • 6 employees • Toronto, ON, Canada
Evidence is an open source, code-based alternative to drag-and-drop BI tools. Build polished data products with just SQL and markdown.
b2b
data-engineering
developer-tools
analytics
data-visualization
Whaly\ \ Y Combinator LogoS2021\ \ • Active • 3 employees • Paris, France
Whaly helps data teams save time on maintenance and analysis building while making business users more autonomous on the analysis they want to improve their decision making.
We do this by providing a self service data platform where both data and business teams can work together.
We understood that most data teams were ending up being a bottleneck for the rest of the company and needed to give more autonomy to business teams to back their decisions with data.
Emilien, Florian and Pierre were the minds behind the Data advertising platforms of the major media and e-commerce companies in France in their earlier position as Product Manager and head of Customer Success, giving them an edge on how to execute successfully a data project.
data-engineering
Whalesync makes data syncing easy. Our automation platform syncs data between key business tools like Webflow/Wix/WordPress and Airtable/Notion/Google Sheets. We give marketing teams two-way, real-time sync, so they can manage their website from their favorite collaboration tools.
Whalesync launched during Y Combinator’s S21 cohort. Since then we’ve raised from some of the world’s top investors. We’re now trusted by hundreds of companies like [Ramp](https://ramp.com/), [Webflow](https://webflow.com/), and [Alchemy](https://www.alchemy.com/), and process millions of transactions every day. Many of our customers enjoy the product so much they [tell all their friends](https://whalesync.com/customers).
data-engineering
no-code
saas
remote-work
web-development
authzed\ \ Y Combinator LogoW2021\ \ • Active • 31 employees • New York City
We build the tools companies need to provide performant and scalable authorization for their applications.
We’re founded by 3 successful entrepreneurs with expertise in enterprise software, most recently as leaders at Red Hat. Jake and Joey met on the APIs team at Google in 2010. They went on to found Quay, where Jimmy joined as their first hire. Over the past decade, they’ve changed the landscape for building and deploying software.
security
data-engineering
developer-tools
open-source
artificial-intelligence
Clear\ \ Y Combinator LogoW2021\ \ • Active • 2 employees • London
On the surface, Clear is a free mobile app where consumers track routines, products and selfies, get brand-agnostic recommendations, and learn from others in the community. Underneath, we’re building a vertically focused AI/data platform that helps consumers, clinicians, and brands understand what actually works, for whom, and why.
The skincare industry is worth over $200B globally, yet it still runs largely on guesswork: consumers waste money on trial-and-error, brands lack real-world evidence, and dermatologists have limited visibility into day-to-day product use and outcomes.
Clear is brand-agnostic, data-rich, and community-driven. Every routine log, product review, and progress update helps create a unique longitudinal dataset across skin concerns, routines, and outcomes - powering better decisions for users and better innovation for the wider industry.
We were the 2022 L’Oréal Beauty Tech for Good winners, have been recognised by Beiersdorf, and have been featured multiple times by Apple, including App of the Day and Best New Apps.
consumer
digital-health
marketplace
data-engineering
Prequel\ \ Y Combinator LogoW2021\ \ • Active • 9 employees • New York City
Prequel makes it easy for companies to share data with their customers. It helps you export data directly to your customer's Snowflake, Redshift, BigQuery, Databricks, or other data warehouse on an ongoing basis.
data-engineering
saas
analytics
Polytomic\ \ Y Combinator LogoW2020\ \ • Active • 7 employees • San Francisco
Polytomic is a no-code web app to sync data between your internal databases, business systems (e.g. Stripe, Salesforce, etc), data warehouses, spreadsheets, and even HTTP APIs.
saas
b2b
data-engineering
Datafold\ \ Y Combinator LogoS2020\ \ • Active • 30 employees • New York City
Datafold automates manual work in data engineering.
We leverage agentic AI to automate both day-to-day tasks, such as testing and code reviews, and massive one-off projects, such as data platform code migrations. Companies from Perplexity to Disney use Datafold to unlock more value from their data by freeing up their data teams from manual work, accelerating developer velocity, and ensuring data quality.
data-engineering
saas
analytics
ai
Mozart Data\ \ Y Combinator LogoS2020\ \ • Active • 24 employees • San Francisco
Mozart Data provides an out-of-the-box modern data stack that empowers anyone to easily consolidate, organize, and prepare their data for analysis. Spin up a data stack that’s built on a best-in-class data warehouse and ETL tool in hours, without any engineering. You can finally spend more time on generating insights and less time wrangling your data.
saas
b2b
data-engineering
Supabase\ \ Y Combinator LogoS2020\ \ • Active • 120 employees • San Francisco
Supabase is the easiest way to get started with Postgres.
Each project within Supabase is an isolated Postgres cluster, allowing customers to scale independently, while still providing the features that you need to build: instant database setup, auth, row level security, realtime data streams, auto-generating APIs, and a simple to use web interface.
We are 100% remote.
open-source
databases
data-engineering
big-data
developer-tools
Dataland\ \ Y Combinator LogoS2020\ \ • Active • 4 employees • New York City
Dataland is the applied AI lab for complex operations and customer support. In 8 months, we've grown to multi-million dollar ARR as just two founders and are highly profitable. We’re partnering with some of the fastest growing startups and public companies in the world.
b2b
data-engineering
data-visualization
ai
Airbyte\ \ Y Combinator LogoW2020\ \ • Active • 90 employees • San Francisco
Airbyte Agents is the context layer your AI agents query before they act. Features the Context Store, which unifies entities across Salesforce, Stripe, Zendesk and more. Use it from the UI, your LLM via MCP, or our SDK. 40% fewer tool calls, up to 80% fewer tokens.
Airbyte Data Replication is the leading open data movement platform that empowers data teams in the AI era by transforming raw data into actionable intelligence.
developer-tools
open-source
data-engineering
ai
artificial-intelligence
Logarithm Labs\ \ Y Combinator LogoW2020\ \ • Active • 2 employees • San Francisco
Easy button to use data for your daily operations. Power your business workflows with quality data.
Logarithm Labs helps you turn manual data wrangling and ad-hoc scripts into repeatable pipelines for your operational workflows. Power your workflows with quality data.
Our product and team of experts do the heavy lifting so that can focus on the business logic that drives your organization.
To learn more, contact us at hello@logarithmlabs.com.
data-engineering
developer-tools
Operator.io\ \ Y Combinator LogoW2020\ \ • Active • 6 employees • New York City
Operator.io is the easiest way to set up an OpenClaw agent and connect it to 1000+ different applications like Gmail, Notion, Stripe, and more. Onboard an AI executive assistant that tackles the tasks you know need to be done that you don't have bandwidth for.
data-engineering
generative-ai
Embrace\ \ Y Combinator LogoS2019\ \ • Active • 62 employees • Los Angeles
*S19 and YCG F21* Engineering teams shouldn’t be in the dark when it comes to the performance of their web and mobile apps and user experiences. With our industry-leading observability technology, Embrace aims to give SRE and DevOps teams better insight into what matters most – their users – by tying performance data to backend observability insights.
As the only user-focused, use-focused observability solution built on OpenTelemetry, Embrace delivers crucial insights across both DevOps and frontend teams to illuminate real customer impact – not just server impact – to deliver the best app experiences. Customers like The New York Times, Marriott, Warby Parker, Masterclass, Home Depot, and Cameo love Embrace’s observability platform because it makes extremely complicated and voluminous data actionable.
open-source
saas
artificial-intelligence
data-engineering
TRM Labs\ \ Y Combinator LogoS2019\ \ • Active • 400 employees • San Francisco
TRM is on a mission to build a safer world for billions of people.
AI is extending the capability gap between attackers and defenders. The fundamental unit of scale has shifted from humans to AI tokens.
TRM combines proprietary data, with an AI investigations platform to build a system of action for crime fighters. Think Claude CoWork for mission critical investigations.
We believe AI scale crime deserves AI scale defense.
Join our mission ➔ www.trmlabs.com/careers
machine-learning
data-engineering
fintech
cybersecurity
govtech
Gecko Robotics\ \ Y Combinator LogoW2016\ \ • Active • 230 employees • Pittsburgh, PA, USA
Gecko Robotics is the pioneer of AI + Robotics [AIR technology], transforming how the world builds, operates, and maintains its most critical infrastructure for a more reliable and sustainable future. Using fixed sensors and robots that climb, crawl, swim, and fly, we combine first-order data layers with the predictive power of AI into a single source of truth for the physical world. Cantilever™ is our operating platform, powered by AIR technology, that empowers teams to achieve operational excellence through actionable data for immediate and long-term planning.
big-data
energy
robotics
data-engineering
artificial-intelligence
Mezmo\ \ Y Combinator LogoW2015\ \ • Active • 172 employees • Miami, FL, USA
Mezmo, formerly LogDNA, is an observability platform to manage and take action on your data. It ingests, processes, and routes log data to fuel enterprise-level application development and delivery, security, and compliance use cases.
Mezmo was brought to life by three-time co-founders Chris Nguyen and Lee Liu and included in the Winter 2015 batch of Y Combinator. In 2018 the company partnered with tech giant, IBM, to become the sole logging provider for IBM Cloud.
Mezmo is on a mission to empower people who build solutions that shape the world. We’re doing this by delivering a platform that enables enterprises to get more value from their observability data in real time, regardless of source, destination, use case, or scale. We’re not the only ones working on this problem but we have a few things the others don’t.
We’re cloud-native and know how to make the most of modern technology like Kubernetes. We have scaled a solution from zero to petabyte scale in a short amount of time, while supporting thousands of active users across multiple environments. We are hungry for change and are surrounded by enterprises telling us they’re hungry, too. We have a kick-ass group of people who are thinking about the problem analytically and are excited to change the observability world for the better. Mezmo has helped some of the world’s most innovative companies transform how they manage their systems and applications. Still, we know that we can help them get more value from their observability data by providing more flexibility and control over how they use it. This will enable teams to spend less time switching between data silos so they can focus on shipping better, more resilient, and secure products.
We have momentum on our side. Last year we saw triple digit revenue growth and added 800 new customers to our roster. Recent accolades include being named to YC’s Top Companies, CRN’s 10 Hottest DevOps Startups, and EMA’s Top 3 Observability Platforms.
data-engineering
devsecops
kubernetes
saas
developer-tools
Etleap\ \ Y Combinator LogoW2013\ \ • Active • 11 employees • San Francisco
Etleap is an ETL solution for creating perfect data pipelines from day one. Unlike other enterprise solutions, Etleap doesn’t require extensive engineering work to set up, maintain, and scale. It automates most ETL setup and maintenance work, and simplifies the rest into 10-minute tasks that analysts can own.
data-engineering
PeerDB\ \ Y Combinator LogoS2023\ \ • Acquired • 2 employees
At PeerDB, we are building a fast, simple and the most cost effective way to stream data from Postgres to Data Warehouses, Queues and Storage engines. If you are running Postgres at the heart of your data-stack and move data at scale from Postgres to any of the above targets, PeerDB can provide value.
We support different modes of streaming - log based (CDC), cursor based (timestamp or integer) and XMIN based. Performance wise, we are 10x faster than existing tools. Features wise, we support native Postgres features such as comprehensive set of data-types incl. jsonb/arrays/postgis, efficiently streaming toast columns, schema changes and so on.
open-source
enterprise-software
data-engineering
databases
developer-tools
Tarsal\ \ Y Combinator LogoS2021\ \ • Acquired • 10 employees • New York City
Tarsal is a data pipeline custom built for security teams. As security data grows 25% year over year, security teams desperately need access to best-in-class data infrastructure. Tarsal bridges the gap between the modern data stack and security teams, pioneering the modern security data stack.
cybersecurity
data-engineering
big-data
b2b
Lume\ \ Y Combinator LogoW2023\ \ • Acquired • 5 employees
Lume speeds up customer implementation with AI. Lume helps teams analyze, map, and ingest customer data up to 87% faster, accelerating time-to-revenue.
data-engineering
b2b
saas
infrastructure
ai
Outerbase\ \ Y Combinator LogoW2023\ \ • Acquired • 4 employees • Pittsburgh, PA, USA
Outerbase is the interface for your database. Companies use Outerbase to view, edit, and modify their data and even generate beautiful visual dashboards without having to write a single line of SQL.
data-engineering
analytics
developer-tools
generative-ai
ai
Versori\ \ Y Combinator LogoW2023\ \ • Acquired • 16 employees • Manchester, UK
Orchestrate custom integrations, workflows & agents in hours, not months. Take control of your integration strategy and breathe easy with maintenance on AI Autopilot.
For Product Teams: Build better integration libraries. Build a feature-rich integration library, for your users to enjoy. Offer out-of-the-box integrations that work for you and your customers. Embedded IPaaS typically locks you into connector or endpoint limitations. Versori gives you to tools for limitless customisation. Proactive, self healing agents, scan your connected apps for endpoint or schema changes. You get alerted, Versori AI fixes the change. Embed Versori built integrations into your app with the Versori SDK. Flexible to your development approach with advanced user management.
For Operations Teams: Get your internal systems speaking the same language. Deliver integrations for new software in days, not months—so you can start unlocking value immediately. Versori’s speed to value reduces typical deployment fees by half—or more. Low code for speed. Full code for control. No more limits from inflexible integration platforms.
For GTM & Sales Teams: Say yes to any prospect's integration request. Stop bouncing between teams to get integrations built. With Versori, Sales can go straight to yes. No more escalations or delays. Versori offer fully managed custom-builds, so your customers get exactly what they need, without compromise.
api
b2b
data-engineering
no-code
saas
Satsuma\ \ Y Combinator LogoS2021\ \ • Acquired • 5 employees • San Francisco
Satsuma is a developer tool for building applications on top of real-time blockchain data. Our product lets developers take decoded data from multiple chains, customize it for their use cases, and access it through API endpoints.
Blockchains serve as distributed databases for these products, holding their most important data. However, it’s difficult to access and query that data. We believe this friction is an enormous blocker for web3 developers and that better tooling will enable mass adoption for web3.
We’re a founding team of engineers, having built data infrastructure and product as early employees at Airtable, Heap, and Y Combinator.
crypto-web3
data-engineering
developer-tools
saas
Stackshine\ \ Y Combinator LogoW2022\ \ • Acquired • 7 employees • Portland, OR, USA
Stackshine is creating mission control for enterprise IT teams. We discover all the software being used across their organization and then automate workflows related to onboarding/offboarding, cost savings, and security.
enterprise
analytics
data-engineering
robotic-process-automation
productivity
HomeRoom\ \ Y Combinator LogoW2022\ \ • Acquired • 25 employees • San Francisco
Homeroom helps investors provide affordable housing while making a 22% ROI.
We do this by sourcing properties, arranging capital, managing construction, vetting tenants and collecting rent by the room. To date, Homeroom has brought on 85 property investors, growing 6X annually, are bringing in 420K in annualized net-revenue
How it works: We help investors buy homes in cities that are attractive to young people, but lack affordable housing options. We then renovate and after about 20 days, the home is ready and we find qualified renters by the room.
We launched in 2018 in Kansas City with 1 home. We now have 105 homes in 31 cities. In 2021, we grew rental GMV to $1.8M (300% YoY growth). Our average rent across every property is $458, which is about 50% lower than market comps, and our investors see returns up to 50% higher.
We are HomeRoom. Johnny is the financial analyst/domain expert. Thomas is a cereal entrepreneur with a PHD in ML, and Mike hacked growth for Airbnb and Facebook.
real-estate
proptech
machine-learning
nlp
data-engineering
Sarus\ \ Y Combinator LogoW2022\ \ • Acquired • 16 employees • Paris, France
Sarus solves the problem of accessing or sharing personal data for analytics or machine learning. The solution deploys natively in data infrastructures and lets practitioners work on data they cannot see. Every interaction with the sensitive data is protected with the highest privacy standard: differential privacy Sarus makes traditional anonymization methods irrelevant, saving months in compliance and data engineering while preserving all of the value of data.
data-engineering
analytics
compliance
Hydra\ \ Y Combinator LogoW2022\ \ • Acquired • 6 employees • San Francisco
Hydra is a real-time analytics database management system for Postgres. We seperate compute from storage to offer software engineers serverless analytics with autoscale, write isolation, automatic caching, and more.
Shipping scalable projects on time series and event data has never been easier. Hydra is available for local development, cloud, and bare metal deployment.
open-source
data-engineering
developer-tools
analytics
Bracket\ \ Y Combinator LogoW2022\ \ • Acquired • 3 employees • New York City
Bracket is the two-way data pipeline between popular business tools and backend databases. When ops teams update data in Salesforce or Airtable, and engineers update data in the database, Bracket connects the two sources to reflect the same information.
saas
b2b
data-engineering
Secoda\ \ Y Combinator LogoS2021\ \ • Active • 27 employees • Toronto, ON, Canada
We believe that finding the right data shouldn’t require a technical background or hours of digging. That’s why Secoda applies AI to transform messy, siloed analytics into a searchable, intuitive knowledge layer—so every question gets a fast, useful answer.
Our vision is to become the AI search engine for your company’s analytics, making data discovery as seamless as finding a website on Google. To do that, Secoda gets data AI-ready by unifying governance, cataloging, observability, and lineage into one trusted, easy-to-use platform—empowering every team to move faster, stay compliant, and make smarter decisions.
data-engineering
analytics
saas
b2b
artificial-intelligence
Avenue\ \ Y Combinator LogoW2021\ \ • Acquired • 8 employees • New York City
Avenue is a simple way for business teams to set up alerts from their database or data warehouse. Think Datadog / PagerDuty for operations teams.
Operations teams create set-and-forget alerts on all their data, so they can be more proactive with their time (and monitor on more nuanced triggers than just what fits on their dashboard page).
Avenue can improve response times to critical problems from several days to real-time by alerting directly on the data sources that customers already use.
data-engineering
saas
developer-tools
Metaplane\ \ Y Combinator LogoW2020\ \ • Acquired • 32 employees • New York City
Metaplane ensures everyone trusts the data that powers your business. Data teams at Bose, Ramp, and Klaviyo use our data observability platform to prevent and detect data issues — before the CEO pings them about weird revenue numbers.
We do this with ML-based anomaly detection, end-to-end column-level lineage, and tools to help prevent incidents before they occur. You can monitor your entire data stack within 30 minutes.
The company is backed by Khosla Ventures, Y Combinator, and the founders of Okta, HubSpot, and Vercel.
data-engineering
developer-tools
saas
Chaos Genius\ \ Y Combinator LogoW2020\ \ • Acquired • 10 employees • Bengaluru
Chaos Genius is a DataOps Observability platform for Snowflake. Enable Snowflake Observability to reduce Snowflake costs and optimize query performance.
cloud-workload-protection
machine-learning
data-engineering
analytics
open-source
Jitsu\ \ Y Combinator LogoS2020\ \ • Acquired • 4 employees • New York City
Jitsu is the fastest, most durable way to collect event data from every source - web, app, email, chatbot, CRM - into your data warehouse. 100% open-source. Purpose built, secure and ready in minutes.
data-engineering
saas
b2b
open-source
Data Mechanics\ \ Y Combinator LogoS2019\ \ • Acquired • 25 employees • Paris, France
Data Mechanics was acquired by NetApp in 2021 and integrated in the Spot.io product portfolio. Our managed Spark-on-Kubernetes platform is live and running under the name Ocean for Apache Spark: https://spot.io/products/ocean-apache-spark/
data-engineering
b2b
saas
open-source
TetraScience\ \ Y Combinator LogoS2015\ \ • Active • 100 employees • Boston
TetraScience provides the world’s first and only R&D Data Cloud, with a mission to transform life sciences R&D, accelerate discovery, and improve human life. Scientists at global pharma and biotech organizations rely on our innovative Tetra Data Platform for easy access to centralized, harmonized, and actionable scientific data to accelerate their digital lab transformation. With best-in-class SaaS performance, a team of industry innovators, and excellent product/market fit, Tetra is positioned to become an iconic life sciences software company.
saas
data-engineering
Yhat (YC W15, pronounced y-hat) was an end-to-end data science platform. Acquired by Alteryx (NYSE:AYX)
data-engineering
machine-learning
enterprise
artificial-intelligence
Scuba\ \ Y Combinator LogoW2013\ \ • Acquired • 51 employees
Scuba is the fast and scalable event-based analytics solution to answer critical business questions about how customers behave and products are used. Interana allows users to analyze and explore the key business metrics that matter most in a data-driven world – such as growth, retention, conversion and engagement – in seconds, rather than the hours or days it often takes with existing solutions. Interana allows customers to discover and investigate these key insights easily through its visual and interactive interface, which makes data analysis a natural extension of everyone’s workflow.
analytics
big-data
data-engineering
data-visualization
BackType\ \ Y Combinator LogoS2008\ \ • Acquired0 • San Francisco
saas
data-engineering
Hottest Startup Categories
Startups by Industry
- AI
- AI Assistant
- AI-Enhanced Learning
- AIOps
- API
- Aerospace
- Agriculture
- Analytics
- Apparel and Cosmetics
- Asset Management
- Automation
- Automotive
- Aviation and Space
- B2B Software and Services
- Banking and Exchange
- Biotech
- Climate
- Cloud Computing
- Community
- Compliance
- Computer Vision
- Construction
- Consumer
- Consumer Electronics
- Consumer Finance
- Consumer Health Services
- Consumer Health and Wellness
- Content
- Conversational AI
- Credit and Lending
- Crypto / Web3
- Cybersecurity
- Data Engineering
- Deep Learning
- Defense
- Delivery
- Design Tools
- DevOps
- Developer Tools
- Diagnostics
- Digital Health
- Drones
- Drug Discovery and Delivery
- E-commerce
- Education
- Energy
- Engineering, Product and Design
- Enterprise
- Enterprise Software
- Entertainment
- Finance
- Finance and Accounting
- Financial Technology and Services
- Food and Beverage
- Gaming
- Generative AI
- GovTech
- Government
- HR Tech
- Hard Tech
- Hardware
- Health & Wellness
- Health Tech
- Healthcare
- Healthcare IT
- Healthcare Services
- Home and Personal
- Housing and Real Estate
- Human Resources
- Industrial Bio
- Industrials
- Infrastructure
- Insurance
- Investing
- Job and Career Services
- Legal
- LegalTech
- Logistics
- Machine Learning
- Manufacturing
- Manufacturing and Robotics
- Marketing
- Marketplace
- Medical Devices
- Neobank
- Office Management
- Open Source
- Operations
- Payments
- Productivity
- Proptech
- Real Estate and Construction
- Recruiting
- Recruiting and Talent
- Retail
- Robotics
- SaaS
- Sales
- Security
- Social
- Supply Chain
- Supply Chain and Logistics
- Therapeutics
- Transportation Services
- Travel, Leisure and Tourism
- Video
- Virtual and Augmented Reality
- Workflow Automation
- eLearning