Data Engineering Startups funded by Y Combinator (YC) 2026 | Y Combinator

Data Engineering Startups funded by Y Combinator (YC) 2026

September 2026

Browse 87 of the top Data Engineering startups funded by Y Combinator.

We also have a Startup Directory where you can search through over 5,000 companies.

Fivetran\ \ Y Combinator LogoW2013\ \ • Active • 1,200 employees • San Francisco

Fivetran automates data movement out of, into and across cloud data platforms. We automate the most time-consuming parts of the ELT process from extracts to schema drift handling to transformations, so data engineers can focus on higher-impact projects with total pipeline peace of mind. With 99.9% uptime and self-healing pipelines, Fivetran enables hundreds of leading brands across the globe, including Autodesk, Conagra Brands, JetBlue, Lionsgate, Morgan Stanley, and Ziff Davis, to accelerate data-driven decisions and drive business growth. Fivetran is headquartered in Oakland, California, with offices around the world.

data-engineering

saas

analytics

b2b

CueBench\ \ Y Combinator LogoS2026\ \ • Active • 3 employees • San Francisco

data-engineering

reinforcement-learning

Magma\ \ Y Combinator LogoS2026\ \ • Active • 1 employees • San Francisco

We enable companies monetize their agents' traces.

ai

reinforcement-learning

data-engineering

Hebbian Robotics\ \ Y Combinator LogoS2026\ \ • Active • 2 employees • San Francisco

Hebbian is an open source SDK for building quality control pipelines for Physical AI. We enable data teams to scale their pipeline from thousands to millions of hours of data, without managing infrastructure.

https://github.com/Hebbian-Robotics/hflow

robotics

data-engineering

data-science

databases

artificial-intelligence

Olam Labs\ \ Y Combinator LogoS2026\ \ • Active • 2 employees • San Francisco

Multi-agent environments let us evaluate and train models in complex, simulated worlds. We work with researchers on both evaluations and training for character and agentic performance typically difficult to assess for using current popular datasets.

Play one of our first releases, Multi-Agent Arena (https://olamlabs.ai/arena), where humans come play social strategy games against multiple AI agents. It's currently top 50 on OpenRouter and has thousands of matches played each day. We use Multi-Agent Arena to build datasets on agentic performance and real-world socialization.

reinforcement-learning

gaming

data-engineering

artificial-intelligence

Alkera AI\ \ Y Combinator LogoS2026\ \ • Active • 3 employees • San Francisco

Alkera is a data engineering/analysis/science agent that works in your IDE or CLI. If you’re tired of agents that produce false, or even dangerous, reports about your data, then Alkera is for you. Our agent is equipped with your data’s history and lineage, your team’s knowledge, and tailored plugins for data apps such as Snowflake and dbt. Simply install and connect your services to begin doing trusted agentic data work. We also offer VPC & on-prem deployment options for when data can't leave your boundary.

data-science

data-engineering

developer-tools

enterprise-software

artificial-intelligence

Litmus\ \ Y Combinator LogoS2026\ \ • Active • 3 employees • New York City

Litmus is building the most accurate framework for evaluating and benchmarking human capability, starting with software.

AI will compound small differences in human capability into increasingly large differences in what people can accomplish, while making existing static benchmarks obsolete. Software is already there: AI can hill-climb any output-based evaluation, while the ability to direct it is becoming the defining advantage.

Litmus applies the same approach we already use for model evals to humans – creating world-like environments, and inspecting trajectory instead of just output.

Every knowledge industry will soon face the same problem. We build Litmus to tell you what humans are capable of.

We're already helping build frontier technical teams at Mercor, Composio, Neo Scholars, and more.

recruiting

data-engineering

ai

Petrarch\ \ Y Combinator LogoS2026\ \ • Active • 3 employees • San Francisco

Petrarch brings internal company data to frontier labs, starting with manufacturing and industrial companies. We source data including codebases, project files, and payments from bankruptcy courts, then de-identify and prepare the data for model training.

b2b

marketplace

data-engineering

Markov\ \ Y Combinator LogoS2026\ \ • Active • 2 employees • San Francisco

Markov sells expert computer-use data to frontier labs. We’ve sold more than 33k+ hours and our open-source datasets have got over 200k+ downloads on hugging face.

reinforcement-learning

data-labeling

data-engineering

Corvera\ \ Y Combinator LogoW2026\ \ • Active • 4 employees • San Francisco

Corvera (YC W26) is the context layer for AI-native CPG brands. We help brands unlock the full potential of AI by making their unified data legible to any AI tool via MCP.

As of mid-March 2026, we scaled from $0 to $33k in MRR in 4 weeks, are serving 12 brands, and are growing 130% week-on-week.

With Corvera in place, brands can enable anyone in their organization to: - Rapidly deploy AI agents via tools such as Claude and ChatGPT, - Build dashboards and apps using Cursor and Lovable, and - Automate workflows across the entire business, from category management to supply chain.

All in a secure, audit-logged, and access-controlled environment, without the need for a team of data engineers or AI specialists, in a simple plug-and-play solution.

We're founded by Chris (2x Founder; ex-CEO at Better Nature; Forbes 30U30), Dirk (ex-Google Data & AI Lead), and Matthew (ex-Head of Product at Rosemark; Princeton MEng in CS).

To-date, we’ve raised $6.2M and are backed by YC, firstminute capital, 6 Degrees Capital, 20VC, Rebel Fund, Duke Capital Partners, and more. Our ICP are omnichannel retail-focused CPG brands at the $100M+ revenue inflection point.

Our mission is bold: to enable the next generation of unicorn CPG brands to be run by teams of less than 10 people.

Our vision is even bolder: to be the system of intelligence for the $5.4T global CPG industry.

Visit www.corvera.ai to learn more!

b2b

saas

data-engineering

consumer

ai

Velum Labs\ \ Y Combinator LogoW2026\ \ • Active • 2 employees • San Francisco

Velum is the operating system for data quality. Velum automatically monitors and enforces data quality across a company's data stack, so bad data never reaches dashboards. We turn data quality from a manual task into infrastructure that runs itself.

Data trust you can prove. From the pipeline to the boardroom.

machine-learning

data-engineering

Captain\ \ Y Combinator LogoW2026\ \ • Active • 2 employees • San Francisco

Captain syncs complex, multimodal files and cloud storage from sources like S3, SharePoint, and Google Drive, and makes that knowledge searchable for your agents.

Instead of stitching together parsers, embeddings, vector databases, rerankers, and retrieval infrastructure, Captain self-tunes the retrieval pipeline to your data and workloads.

On benchmarks, this improves accuracy from roughly 78% with standard RAG to over 97% by tuning retrieval to the unique structure and content of your data.

Captain is purpose-built for complex, regulated, and high-stakes file search workloads.

Learn more at https://captain.dev

data-engineering

infrastructure

b2b

api

Velvet\ \ Y Combinator LogoF2025\ \ • Active • 4 employees • San Francisco

Infra and data for interactive AI.

ai

data-engineering

conversational-ai

generative-ai

Hyperspell\ \ Y Combinator LogoF2025\ \ • Active • 8 employees • San Francisco

Hyperspell is your Company Brain. AI agents are brilliant and clueless. They ace any test and still have no idea how your company works. Hyperspell connects your tools and synthesizes documents and conversations into a live, permissioned context graph. Any agent can read from it and write back to it like a filesystem, with every fact traced to source.

ai

machine-learning

saas

data-engineering

DeepAware AI (Robotics Center of Silicon Valley)\ \ Y Combinator LogoS2025\ \ • Active • 4 employees • San Francisco

DeepAware (Robotics Center of Silicon Valley, https://roboticscenter.ai/) is the fastest way for enterprises and researchers to get robots and robotics parts in the US — 72-hour delivery or Bay Area pickup.

Beyond hardware, we help teams collect teleoperation data, build reinforcement learning environments, and deploy robots into production. Customers include AI labs, industrial operations, research teams, and event producers.

supply-chain

data-engineering

robotics

machine-learning

artificial-intelligence

sieve\ \ Y Combinator LogoP2025\ \ • Active • 2 employees • New York City

sieve solves data cleaning for hedge funds and investment firms by letting them get clean data in four lines of code. Currently, their data pipelines have conditions that raise for human review, which literally send an email to engineers with data that needs to be reviewed. We provide an API that integrates directly into their existing pipeline - instead of raising for human review, they can send all the same information to our API and get clean, high-quality data back.

By using our AI agents built specifically for financial data collection, along with expert-in-the-loop review, we provide our clients with clean, validated data at a scale and level of quality that wasn't achievable before.

apis

investing

data-engineering

Vision Lab\ \ Y Combinator LogoP2025\ \ • Active • 11 employees • San Francisco

We capture and structure real factory workflows at scale by combining first-person industrial video with SOP-level process knowledge.

This enables robotics and AI labs to train on real production data, not just controlled lab environments.

manufacturing

ai

robotics

data-engineering

Nitrode\ \ Y Combinator LogoW2025\ \ • Active • 10 employees • San Francisco

Nitrode is a research company focused on improving game development with AI. To advance game development, we believe AI models firstly need to be judged against robust evaluation frameworks that reflect the actual complexity of the field.

artificial-intelligence

data-engineering

machine-learning

b2b

Kilvin\ \ Y Combinator LogoF2024\ \ • Active • 3 employees • New York City

Kilvin is an AI-native agency that builds custom software solutions for non-technical teams.

With AI, the future of software is democratized. Every team deserves tools built exactly for them, not off-the-shelf SaaS products they need to adapt to. Kilvin makes it a reality, delivering a private, secure workspace of custom applications, built around your workflows and ready in days, not months.

No engineers needed on your side. Just software that fits exactly how you work.

developer-tools

automation

ai

data-engineering

productivity

Melder\ \ Y Combinator LogoF2024\ \ • Active • 2 employees • New York City

Melder is an Excel add-in that brings AI functions and document support into your spreadsheets. Upload files directly into cells, use smart formulas like =GEN, and build automations—all without leaving Excel.

Core features: - File-to-Sheet: Drop PDFs directly into cells, then reference them in formulas. - AI-Powered Functions: Write formulas like =GEN() or =EXTRACT() to summarize, classify, and analyze content. - Chat Assistant: Use our AI assistant to help build your sheets or answer questions from your data, live in the workbook.

Business users use Melder to: - Accelerate diligence by extracting insights from data rooms - Review contracts by identifying key terms and clauses instantly - Run market research by pulling information from competitor websites - Synthesize transcripts by generating summaries from interviews and calls

Melder brings the power of structured spreadsheet logic to the messy, unstructured data world—no coding needed.

data-engineering

generative-ai

artificial-intelligence

Sensei\ \ Y Combinator LogoS2024\ \ • Active • 2 employees • Boston

Sensei helps robotics companies scale and outsource their training data collection. Our hardware platform enables the collection of human-demonstration data at a tenth of the cost and twice the speed of current teleop approaches. Our software platform acts like Scale AI for robotics data: a large network of paid human operators use our low-cost collection platform to fulfill data-generation requests.

robotics

hard-tech

artificial-intelligence

marketplace

data-engineering

Overstand Labs\ \ Y Combinator LogoW2025\ \ • Active • 4 employees • New York City

Overstand is a data lab that allows our customers to navigate any set of data in just a few minutes.

*Enterprise*: For enterprises, we we unify Slack, email, calls, and operational data, then surface the signals that matter — customer needs, risks, and revenue opportunities hidden in everyday conversations.

*Legal Firms*: For legal firms, we help them really quickly understand their entire discovery corpus (either before, or after document review), and quickly build out an initial case assessment and facts.

Instead of waiting on reports or manual analysis, teams get immediate, evidence-backed answers from the data they already have. Overstand delivers clarity and leverage as your business scales.

data-science

data-engineering

conversational-ai

artificial-intelligence

legaltech

Sharpe\ \ Y Combinator LogoS2024\ \ • Active • 3 employees • San Francisco

Sharpe helps traders go from idea to profit in minutes with AI, bundling petabytes of market data with high-performance infrastructure.

ai

finance

analytics

data-engineering

big-data

Zeit AI\ \ Y Combinator LogoS2024\ \ • Active • 12 employees • Munich, Germany

Ask a question in plain English and ZeitMind connects SAP, your ERP, CRM, HR systems and 600+ other sources, prepares the messy data, and turns it into the live applications your teams actually make decisions with.

It is built for the finance, controlling, procurement, supply chain and sales teams at companies that run enterprise systems without an enterprise data team.

artificial-intelligence

analytics

b2b

saas

data-engineering

Mica AI\ \ Y Combinator LogoS2024\ \ • Active • 3 employees • San Francisco

Mica's AI agents replace the data ops teams fixing bad data.

When bad or missing data breaks the pipeline, and orchestration, retries, and monitoring fail, painful manual review work kicks in, pulling humans in to investigate and patch data issues across systems. Mica does what those humans do: gathering the right information from internal docs and external systems, reasoning across context, and resolving errors autonomously to get the pipeline moving again.

The result: dramatically reduce time, cost, and operational drag as your data pipelines scale without scaling ops headcount. Mica turns judgment-heavy data fixes from a manual bottleneck into an automated background process with full auditability.

data-engineering

data-science

enterprise-software

Trellis AI\ \ Y Combinator LogoW2024\ \ • Active • 34 employees • San Francisco

b2b

data-engineering

databases

infrastructure

ai

kater.ai\ \ Y Combinator LogoW2024\ \ • Active • 3 employees • San Francisco

1. You explain your problem. 2. Kater identifies the most important data questions to ask. 3. Kater writes the code. 4. You get insights in seconds rather than weeks.

Kater.ai flips the script on enterprise analytics by making every user an expert analyst. It uses a continuous classification engine to turn a single business question into a contextualized package of questions that is specific to your needs.

Kater puts the power of data into the hands of business experts while ensuring they use trusted data that is specific to their persona. No more waiting for data analysts. No more wasted time on analysis misfires and rework.

Yvonne was a data engineer and analyst who built the entire data stack at CREXi. Robin led engineering in Microsoft.

Data is the new oil. Companies are data-rich, insight-poor. We're helping companies become insight-rich. This is the future of data.

data-engineering

analytics

artificial-intelligence

Ocular AI\ \ Y Combinator LogoW2024\ \ • Active • 8 employees • San Francisco

ai

data-engineering

machine-learning

speech-recognition

Reducto\ \ Y Combinator LogoW2024\ \ • Active • 30 employees • San Francisco

Reducto is the agentic document platform for leading AI teams building systems to automate document workflows at enterprise scale.

Our platform provides a comprehensive toolkit for working with documents the way a human would, combining custom in-house and leading frontier models to power efficient and accurate document workflows. Reducto is trusted by leading AI teams at companies like Harvey, Scale AI, and Vanta.

We are built for enterprise workloads with flexible deployment options from the cloud to fully air-gapped environments, SOC II and HIPAA compliance, and zero data retention.

Learn more: https://reducto.ai/ Find our YC deal: https://reducto.ai/yc

documents

data-engineering

search

enterprise-software

ai

DataShare\ \ Y Combinator LogoS2023\ \ • Active • 1 employees • Austin, TX, USA

DataShare is a data-as-a-service platform that lets you embed charts, dashboards and exports directly into your product. For example, if you run an accounting startup, DataShare would enable you to embed a full profit and loss dashboard, with downloadable statements. DataShare is backed by an enterprise-grade data warehouse, and can be implemented in fewer than 20 lines of code.

data-engineering

databases

analytics

Cedalio\ \ Y Combinator LogoS2023\ \ • Active • 6 employees • San Francisco

Cedalio is the AI agent platform for the Office of the CFO. We automate procurement, AP, and the data flows behind them — from invoice intake to 3-way match to payment — so finance teams stop doing the work and start governing it.

data-engineering

ai

finance

fintech

Artie\ \ Y Combinator LogoS2023\ \ • Active • 16 employees • San Francisco

Artie is software that streams data from databases to data warehouses in real-time. Today, most companies run their ETL process every few hours or overnight, so their data warehouse is always out of date; with Artie, the warehouse always has live production data.

data-engineering

developer-tools

enterprise-software

saas

Ohm\ \ Y Combinator LogoW2023\ \ • Active • 6 employees • San Francisco

Ohm’s AI platform helps Fortune 100 engineering teams accelerate the engineering and testing of complex hardware, including wearables, electric vehicles, and batteries. Launched in 2025, Ohm's platform is used by leading teams to compress development cycles, optimize test programs, and launch better products, faster.

data-engineering

ai

artificial-intelligence

hardware

TableFlow\ \ Y Combinator LogoW2023\ \ • Active • 2 employees • San Francisco

TableFlow builds AI teammates for data tasks, helping operations and data teams automate the messy, manual tasks buried in PDFs, spreadsheets, images, and emails.

ai

automation

documents

data-engineering

saas

Honeydew\ \ Y Combinator LogoW2023\ \ • Active • 6 employees • Tel Aviv-Yafo, Israel

The way people use data is constantly changing. Data teams must support every new context without breaking the shared truth. Honeydew’s semantic layer does it automatically. We validate each change and update every data flow.

Using Honeydew, data teams can support 10x more data users - without more engineers or compromising integrity.

saas

data-engineering

analytics

b2b

Sunpia\ \ Y Combinator LogoS2022\ \ • Active • 3 employees • San Francisco

Sunpia lets developers easily experience the cost and speed benefits of serverless infrastructure, without having to rewrite their code. Developers annotate their code and Sunpia automatically designs a microservice version of it they can deploy on their own cloud.

data-engineering

developer-tools

kubernetes

Midplane\ \ Y Combinator LogoS2022\ \ • Active • 2 employees • Berlin, Germany

Companies are connecting AI coding agents like Cursor and Claude directly to their databases — where one wrong command can delete data or leak customer records. Midplane sits in between and controls what each agent is allowed to do: it blocks dangerous queries, hides sensitive data, and keeps a record of everything the agent touches. It’s open source, and runs either as a hosted service or on your own servers.

b2b

productivity

generative-ai

data-engineering

ai

MovingLake\ \ Y Combinator LogoS2022\ \ • Active • 3 employees • Mexico City, CDMX, Mexico

MovingLake is Fivetran for event-driven architectures. Companies such as Casai use our product to obtain orders and price changes in real time.

analytics

api

data-engineering

b2b

saas

Lamin\ \ Y Combinator LogoS2022\ \ • Active • 10 employees

Improve your workflows with open-source context and data access.

Query, trace, and govern with a lineage-native, format-agnostic lakehouse for agents and teams.

With support for biological formats and registries — by the creators of Scanpy

biotech

data-engineering

machine-learning

developer-tools

open-source

Sanifu\ \ Y Combinator LogoS2022\ \ • Active • 5 employees • Nairobi, Kenya

Sanifu automates the repetitive work spreadsheets are used for. Teams across finance, marketing, operations, and revenue simply describe their repetitive work like cleaning and validating data, matching records, preparing reports and moving data from emails, PDFs and Excel exports into dashboards and other systems. Sanifu captures the steps, rules and outputs so the work is repeated reliably every time

Some of these work never needed spreadsheets to begin with; the likes of data cleanup, bank reconciliations, report consolidation, payment allocations, channel attributions. Spreadsheets have become the standard workbench for handling these tasks and in the process, teams lose valuable time doing manual data entry, filling templates, formatting data, applying formulas, deleting rows or columns.

Sanifu is replacing the spreadsheet while maintaining the logic and processes behind the work, getting you from inputs to your desired outputs in seconds. This significantly reduces time spent on copy-paste work, spreadsheet cleanup, and manual entries and checking, while improving accuracy, consistency, and operational visibility. Sanifu helps teams scale critical back-office processes without relying on brittle spreadsheets, scripts, or general AI tools that reinvent the logic every run.

Sanifu is a full operational platform with 5k+ app integrations, 1k+ native functions, forms, scheduler, data stores, dashboards, approvals, reviews, AI-powered entity matching among many other features that make it an end-to-end data app management platform

b2b

saas

data-engineering

analytics

ai

LanceDB\ \ Y Combinator LogoW2022\ \ • Active • 35 employees • San Francisco

LanceDB is a new open-source vector database that can support low-latency billion-scale vector search on a single node. Built around a new columnar data format, LanceDB makes it incredibly easy to build applications for generative AI, recsys, search engines, content moderation, and more.

aiops

data-engineering

machine-learning

open-source

Elementary\ \ Y Combinator LogoW2022\ \ • Active • 12 employees • Tel Aviv-Yafo, Israel

Elementary enables data teams to detect problems in their data before their users do. An open-source solution that any data engineer can deploy in minutes without sharing sensitive data.

data-engineering

analytics

developer-tools

open-source

Dynamo AI\ \ Y Combinator LogoW2022\ \ • Active • 40 employees • San Francisco

End-to-end privacy, security, and compliance solutions to prepare your organization for emerging AI regulations.

privacy

data-engineering

machine-learning

Sieve\ \ Y Combinator LogoW2022\ \ • Active • 34 employees • San Francisco

Sieve builds the data and environments frontier AI labs use to train the next generation of multimodal systems.

AI is moving beyond chatbots into video, audio, images, software, robotics, and interactive worlds. The next generation of models will need to understand how the world looks, sounds, moves, responds, and changes over time. Progress is bottlenecked by one thing: high-quality data.

Sieve brings together exabyte-scale infrastructure, novel multimodal understanding techniques, large-scale sourcing, and deep research partnerships to create datasets and environments with unmatched precision, quality, and speed. This has earned the trust of frontier AI labs, Fortune 100 companies, and fast-growing AI startups working on generative media, robotics, computer use, world models, and agentic systems.

video

developer-tools

data-engineering

data-labeling

ai

Versable\ \ Y Combinator LogoW2022\ \ • Active • 3 employees • San Francisco

Auto parts retailers get product data from hundreds of manufacturers that is inaccurate and inconsistent, often with big gaps in key values. They currently have a team of "catalog managers" who are required to process and enhance this data line by line, resulting in a week to months long lag between receiving product data and actually being able to start generating revenue from those products.

Versable leverages AI to scan the web for tens of millions of auto parts listings, and uses a fine-tuned LLM with RAG to instantly process, enhance, and transform data. With just a part number, Versable is able to generate market-ready titles, product descriptions, and specs, in any format that's needed.

automotive

manufacturing

data-engineering

ai

Trackingplan\ \ Y Combinator LogoW2022\ \ • Active • 17 employees • Barcelona, Spain

Trackingplan automatically discovers and monitors all the information your applications and websites are collecting, ensuring that you can trust your BI, analytics, marketing, and sales tools.

You can think of us as Segment Protocols but totally transparent, where developers can keep using Google Analytics, Amplitude, Hubspot, Intercom, Braze, etc. as they are used to.

Installed in minutes in using your Tag Manager or adding just one line of code to your web or apps, we model all the data being sent to third parties. Since Trackingplan understands what each piece of data means, it identifies patterns, detects anomalies, and automatically connects the dots to create value from data that was hidden in plain sight:

- An always up-to-date single source of truth and data governance tool. To discover, understand and document your data and improve communication across teams. - Automated notifications when something breaks or changes. To make sure that integrations are always well implemented: Schema errors, traffic anomalies, rogue events... - Easy to understand, customizable, cross-service alerts. To detect trends, insights, and problems without using complex, engineer-oriented solutions.

data-engineering

analytics

saas

Pipekit\ \ Y Combinator LogoS2021\ \ • Active • 9 employees

Our app manages Argo Workflows for data teams, enabling complex data & CI pipelines in half the time while saving companies hundreds of thousands of dollars annually. We maintain Argo Workflows, an open-source pipeline framework for Kubernetes that’s used in production by Bloomberg, Intuit, Adobe, New Relic, NVIDIA, and many other open-source early adopters.

open-source

developer-tools

data-engineering

devops

Evidence\ \ Y Combinator LogoS2021\ \ • Active • 6 employees • Toronto, ON, Canada

Evidence is an open source, code-based alternative to drag-and-drop BI tools. Build polished data products with just SQL and markdown.

b2b

data-engineering

developer-tools

analytics

data-visualization

Whaly\ \ Y Combinator LogoS2021\ \ • Active • 3 employees • Paris, France

Whaly helps data teams save time on maintenance and analysis building while making business users more autonomous on the analysis they want to improve their decision making.

We do this by providing a self service data platform where both data and business teams can work together.

We understood that most data teams were ending up being a bottleneck for the rest of the company and needed to give more autonomy to business teams to back their decisions with data.

Emilien, Florian and Pierre were the minds behind the Data advertising platforms of the major media and e-commerce companies in France in their earlier position as Product Manager and head of Customer Success, giving them an edge on how to execute successfully a data project.

data-engineering

Whalesync\ \ Y Combinator LogoS2021\ \ • Active • 8 employees • 450 8th Ave SE, St. Petersburg, FL 33701, USA

Whalesync makes data syncing easy. Our automation platform syncs data between key business tools like Webflow/Wix/WordPress and Airtable/Notion/Google Sheets. We give marketing teams two-way, real-time sync, so they can manage their website from their favorite collaboration tools.

Whalesync launched during Y Combinator’s S21 cohort. Since then we’ve raised from some of the world’s top investors. We’re now trusted by hundreds of companies like [Ramp](https://ramp.com/), [Webflow](https://webflow.com/), and [Alchemy](https://www.alchemy.com/), and process millions of transactions every day. Many of our customers enjoy the product so much they [tell all their friends](https://whalesync.com/customers).

data-engineering

no-code

saas

remote-work

web-development

authzed\ \ Y Combinator LogoW2021\ \ • Active • 31 employees • New York City

We build the tools companies need to provide performant and scalable authorization for their applications.

We’re founded by 3 successful entrepreneurs with expertise in enterprise software, most recently as leaders at Red Hat. Jake and Joey met on the APIs team at Google in 2010. They went on to found Quay, where Jimmy joined as their first hire. Over the past decade, they’ve changed the landscape for building and deploying software.

security

data-engineering

developer-tools

open-source

artificial-intelligence

Clear\ \ Y Combinator LogoW2021\ \ • Active • 2 employees • London

On the surface, Clear is a free mobile app where consumers track routines, products and selfies, get brand-agnostic recommendations, and learn from others in the community. Underneath, we’re building a vertically focused AI/data platform that helps consumers, clinicians, and brands understand what actually works, for whom, and why.

The skincare industry is worth over $200B globally, yet it still runs largely on guesswork: consumers waste money on trial-and-error, brands lack real-world evidence, and dermatologists have limited visibility into day-to-day product use and outcomes.

Clear is brand-agnostic, data-rich, and community-driven. Every routine log, product review, and progress update helps create a unique longitudinal dataset across skin concerns, routines, and outcomes - powering better decisions for users and better innovation for the wider industry.

We were the 2022 L’Oréal Beauty Tech for Good winners, have been recognised by Beiersdorf, and have been featured multiple times by Apple, including App of the Day and Best New Apps.

consumer

digital-health

marketplace

data-engineering

Prequel\ \ Y Combinator LogoW2021\ \ • Active • 9 employees • New York City

Prequel makes it easy for companies to share data with their customers. It helps you export data directly to your customer's Snowflake, Redshift, BigQuery, Databricks, or other data warehouse on an ongoing basis.

data-engineering

saas

analytics

Polytomic\ \ Y Combinator LogoW2020\ \ • Active • 7 employees • San Francisco

Polytomic is a no-code web app to sync data between your internal databases, business systems (e.g. Stripe, Salesforce, etc), data warehouses, spreadsheets, and even HTTP APIs.

saas

b2b

data-engineering

Datafold\ \ Y Combinator LogoS2020\ \ • Active • 30 employees • New York City

Datafold automates manual work in data engineering.

We leverage agentic AI to automate both day-to-day tasks, such as testing and code reviews, and massive one-off projects, such as data platform code migrations. Companies from Perplexity to Disney use Datafold to unlock more value from their data by freeing up their data teams from manual work, accelerating developer velocity, and ensuring data quality.

data-engineering

saas

analytics

ai

Mozart Data\ \ Y Combinator LogoS2020\ \ • Active • 24 employees • San Francisco

Mozart Data provides an out-of-the-box modern data stack that empowers anyone to easily consolidate, organize, and prepare their data for analysis. Spin up a data stack that’s built on a best-in-class data warehouse and ETL tool in hours, without any engineering. You can finally spend more time on generating insights and less time wrangling your data.

saas

b2b

data-engineering

Supabase\ \ Y Combinator LogoS2020\ \ • Active • 120 employees • San Francisco

Supabase is the easiest way to get started with Postgres.

Each project within Supabase is an isolated Postgres cluster, allowing customers to scale independently, while still providing the features that you need to build: instant database setup, auth, row level security, realtime data streams, auto-generating APIs, and a simple to use web interface.

We are 100% remote.

open-source

databases

data-engineering

big-data

developer-tools

Dataland\ \ Y Combinator LogoS2020\ \ • Active • 4 employees • New York City

Dataland is the applied AI lab for complex operations and customer support. In 8 months, we've grown to multi-million dollar ARR as just two founders and are highly profitable. We’re partnering with some of the fastest growing startups and public companies in the world.

b2b

data-engineering

data-visualization

ai

Airbyte\ \ Y Combinator LogoW2020\ \ • Active • 90 employees • San Francisco

Airbyte Agents is the context layer your AI agents query before they act. Features the Context Store, which unifies entities across Salesforce, Stripe, Zendesk and more. Use it from the UI, your LLM via MCP, or our SDK. 40% fewer tool calls, up to 80% fewer tokens.

Airbyte Data Replication is the leading open data movement platform that empowers data teams in the AI era by transforming raw data into actionable intelligence.

https://github.com/airbytehq/

developer-tools

open-source

data-engineering

ai

artificial-intelligence

Logarithm Labs\ \ Y Combinator LogoW2020\ \ • Active • 2 employees • San Francisco

Easy button to use data for your daily operations. Power your business workflows with quality data.

Logarithm Labs helps you turn manual data wrangling and ad-hoc scripts into repeatable pipelines for your operational workflows. Power your workflows with quality data.

Our product and team of experts do the heavy lifting so that can focus on the business logic that drives your organization.

To learn more, contact us at hello@logarithmlabs.com.

data-engineering

developer-tools

Operator.io\ \ Y Combinator LogoW2020\ \ • Active • 6 employees • New York City

Operator.io is the easiest way to set up an OpenClaw agent and connect it to 1000+ different applications like Gmail, Notion, Stripe, and more. Onboard an AI executive assistant that tackles the tasks you know need to be done that you don't have bandwidth for.

data-engineering

generative-ai

Embrace\ \ Y Combinator LogoS2019\ \ • Active • 62 employees • Los Angeles

*S19 and YCG F21* Engineering teams shouldn’t be in the dark when it comes to the performance of their web and mobile apps and user experiences. With our industry-leading observability technology, Embrace aims to give SRE and DevOps teams better insight into what matters most – their users – by tying performance data to backend observability insights.

As the only user-focused, use-focused observability solution built on OpenTelemetry, Embrace delivers crucial insights across both DevOps and frontend teams to illuminate real customer impact – not just server impact – to deliver the best app experiences. Customers like The New York Times, Marriott, Warby Parker, Masterclass, Home Depot, and Cameo love Embrace’s observability platform because it makes extremely complicated and voluminous data actionable.

open-source

saas

artificial-intelligence

data-engineering

TRM Labs\ \ Y Combinator LogoS2019\ \ • Active • 400 employees • San Francisco

TRM is on a mission to build a safer world for billions of people.

AI is extending the capability gap between attackers and defenders. The fundamental unit of scale has shifted from humans to AI tokens.

TRM combines proprietary data, with an AI investigations platform to build a system of action for crime fighters. Think Claude CoWork for mission critical investigations.

We believe AI scale crime deserves AI scale defense.

Join our mission ➔ www.trmlabs.com/careers

machine-learning

data-engineering

fintech

cybersecurity

govtech

Gecko Robotics\ \ Y Combinator LogoW2016\ \ • Active • 230 employees • Pittsburgh, PA, USA

Gecko Robotics is the pioneer of AI + Robotics [AIR technology], transforming how the world builds, operates, and maintains its most critical infrastructure for a more reliable and sustainable future. Using fixed sensors and robots that climb, crawl, swim, and fly, we combine first-order data layers with the predictive power of AI into a single source of truth for the physical world. Cantilever™ is our operating platform, powered by AIR technology, that empowers teams to achieve operational excellence through actionable data for immediate and long-term planning.

big-data

energy

robotics

data-engineering

artificial-intelligence

Mezmo\ \ Y Combinator LogoW2015\ \ • Active • 172 employees • Miami, FL, USA

Mezmo, formerly LogDNA, is an observability platform to manage and take action on your data. It ingests, processes, and routes log data to fuel enterprise-level application development and delivery, security, and compliance use cases.

Mezmo was brought to life by three-time co-founders Chris Nguyen and Lee Liu and included in the Winter 2015 batch of Y Combinator. In 2018 the company partnered with tech giant, IBM, to become the sole logging provider for IBM Cloud.

Mezmo is on a mission to empower people who build solutions that shape the world. We’re doing this by delivering a platform that enables enterprises to get more value from their observability data in real time, regardless of source, destination, use case, or scale. We’re not the only ones working on this problem but we have a few things the others don’t.

We’re cloud-native and know how to make the most of modern technology like Kubernetes. We have scaled a solution from zero to petabyte scale in a short amount of time, while supporting thousands of active users across multiple environments. We are hungry for change and are surrounded by enterprises telling us they’re hungry, too. We have a kick-ass group of people who are thinking about the problem analytically and are excited to change the observability world for the better. Mezmo has helped some of the world’s most innovative companies transform how they manage their systems and applications. Still, we know that we can help them get more value from their observability data by providing more flexibility and control over how they use it. This will enable teams to spend less time switching between data silos so they can focus on shipping better, more resilient, and secure products.

We have momentum on our side. Last year we saw triple digit revenue growth and added 800 new customers to our roster. Recent accolades include being named to YC’s Top Companies, CRN’s 10 Hottest DevOps Startups, and EMA’s Top 3 Observability Platforms.

data-engineering

devsecops

kubernetes

saas

developer-tools

Etleap\ \ Y Combinator LogoW2013\ \ • Active • 11 employees • San Francisco

Etleap is an ETL solution for creating perfect data pipelines from day one. Unlike other enterprise solutions, Etleap doesn’t require extensive engineering work to set up, maintain, and scale. It automates most ETL setup and maintenance work, and simplifies the rest into 10-minute tasks that analysts can own.

data-engineering

PeerDB\ \ Y Combinator LogoS2023\ \ • Acquired • 2 employees

At PeerDB, we are building a fast, simple and the most cost effective way to stream data from Postgres to Data Warehouses, Queues and Storage engines. If you are running Postgres at the heart of your data-stack and move data at scale from Postgres to any of the above targets, PeerDB can provide value.

We support different modes of streaming - log based (CDC), cursor based (timestamp or integer) and XMIN based. Performance wise, we are 10x faster than existing tools. Features wise, we support native Postgres features such as comprehensive set of data-types incl. jsonb/arrays/postgis, efficiently streaming toast columns, schema changes and so on.

open-source

enterprise-software

data-engineering

databases

developer-tools

Tarsal\ \ Y Combinator LogoS2021\ \ • Acquired • 10 employees • New York City

Tarsal is a data pipeline custom built for security teams. As security data grows 25% year over year, security teams desperately need access to best-in-class data infrastructure. Tarsal bridges the gap between the modern data stack and security teams, pioneering the modern security data stack.

cybersecurity

data-engineering

big-data

b2b

Lume\ \ Y Combinator LogoW2023\ \ • Acquired • 5 employees

Lume speeds up customer implementation with AI. Lume helps teams analyze, map, and ingest customer data up to 87% faster, accelerating time-to-revenue.

data-engineering

b2b

saas

infrastructure

ai

Outerbase\ \ Y Combinator LogoW2023\ \ • Acquired • 4 employees • Pittsburgh, PA, USA

Outerbase is the interface for your database. Companies use Outerbase to view, edit, and modify their data and even generate beautiful visual dashboards without having to write a single line of SQL.

data-engineering

analytics

developer-tools

generative-ai

ai

Versori\ \ Y Combinator LogoW2023\ \ • Acquired • 16 employees • Manchester, UK

Orchestrate custom integrations, workflows & agents in hours, not months. Take control of your integration strategy and breathe easy with maintenance on AI Autopilot.

For Product Teams: Build better integration libraries. Build a feature-rich integration library, for your users to enjoy. Offer out-of-the-box integrations that work for you and your customers. Embedded IPaaS typically locks you into connector or endpoint limitations. Versori gives you to tools for limitless customisation. Proactive, self healing agents, scan your connected apps for endpoint or schema changes. You get alerted, Versori AI fixes the change. Embed Versori built integrations into your app with the Versori SDK. Flexible to your development approach with advanced user management.

For Operations Teams: Get your internal systems speaking the same language. Deliver integrations for new software in days, not months—so you can start unlocking value immediately. Versori’s speed to value reduces typical deployment fees by half—or more. Low code for speed. Full code for control. No more limits from inflexible integration platforms.

For GTM & Sales Teams: Say yes to any prospect's integration request. Stop bouncing between teams to get integrations built. With Versori, Sales can go straight to yes. No more escalations or delays. Versori offer fully managed custom-builds, so your customers get exactly what they need, without compromise.

api

b2b

data-engineering

no-code

saas

Satsuma\ \ Y Combinator LogoS2021\ \ • Acquired • 5 employees • San Francisco

Satsuma is a developer tool for building applications on top of real-time blockchain data. Our product lets developers take decoded data from multiple chains, customize it for their use cases, and access it through API endpoints.

Blockchains serve as distributed databases for these products, holding their most important data. However, it’s difficult to access and query that data. We believe this friction is an enormous blocker for web3 developers and that better tooling will enable mass adoption for web3.

We’re a founding team of engineers, having built data infrastructure and product as early employees at Airtable, Heap, and Y Combinator.

crypto-web3

data-engineering

developer-tools

saas

Stackshine\ \ Y Combinator LogoW2022\ \ • Acquired • 7 employees • Portland, OR, USA

Stackshine is creating mission control for enterprise IT teams. We discover all the software being used across their organization and then automate workflows related to onboarding/offboarding, cost savings, and security.

enterprise

analytics

data-engineering

robotic-process-automation

productivity

HomeRoom\ \ Y Combinator LogoW2022\ \ • Acquired • 25 employees • San Francisco

Homeroom helps investors provide affordable housing while making a 22% ROI.

We do this by sourcing properties, arranging capital, managing construction, vetting tenants and collecting rent by the room. To date, Homeroom has brought on 85 property investors, growing 6X annually, are bringing in 420K in annualized net-revenue

How it works: We help investors buy homes in cities that are attractive to young people, but lack affordable housing options. We then renovate and after about 20 days, the home is ready and we find qualified renters by the room.

We launched in 2018 in Kansas City with 1 home. We now have 105 homes in 31 cities. In 2021, we grew rental GMV to $1.8M (300% YoY growth). Our average rent across every property is $458, which is about 50% lower than market comps, and our investors see returns up to 50% higher.

We are HomeRoom. Johnny is the financial analyst/domain expert. Thomas is a cereal entrepreneur with a PHD in ML, and Mike hacked growth for Airbnb and Facebook.

real-estate

proptech

machine-learning

nlp

data-engineering

Sarus\ \ Y Combinator LogoW2022\ \ • Acquired • 16 employees • Paris, France

Sarus solves the problem of accessing or sharing personal data for analytics or machine learning. The solution deploys natively in data infrastructures and lets practitioners work on data they cannot see. Every interaction with the sensitive data is protected with the highest privacy standard: differential privacy Sarus makes traditional anonymization methods irrelevant, saving months in compliance and data engineering while preserving all of the value of data.

data-engineering

analytics

compliance

Hydra\ \ Y Combinator LogoW2022\ \ • Acquired • 6 employees • San Francisco

Hydra is a real-time analytics database management system for Postgres. We seperate compute from storage to offer software engineers serverless analytics with autoscale, write isolation, automatic caching, and more.

Shipping scalable projects on time series and event data has never been easier. Hydra is available for local development, cloud, and bare metal deployment.

open-source

data-engineering

developer-tools

analytics

Bracket\ \ Y Combinator LogoW2022\ \ • Acquired • 3 employees • New York City

Bracket is the two-way data pipeline between popular business tools and backend databases. When ops teams update data in Salesforce or Airtable, and engineers update data in the database, Bracket connects the two sources to reflect the same information.

saas

b2b

data-engineering

Secoda\ \ Y Combinator LogoS2021\ \ • Active • 27 employees • Toronto, ON, Canada

We believe that finding the right data shouldn’t require a technical background or hours of digging. That’s why Secoda applies AI to transform messy, siloed analytics into a searchable, intuitive knowledge layer—so every question gets a fast, useful answer.

Our vision is to become the AI search engine for your company’s analytics, making data discovery as seamless as finding a website on Google. To do that, Secoda gets data AI-ready by unifying governance, cataloging, observability, and lineage into one trusted, easy-to-use platform—empowering every team to move faster, stay compliant, and make smarter decisions.

data-engineering

analytics

saas

b2b

artificial-intelligence

Avenue\ \ Y Combinator LogoW2021\ \ • Acquired • 8 employees • New York City

Avenue is a simple way for business teams to set up alerts from their database or data warehouse. Think Datadog / PagerDuty for operations teams.

Operations teams create set-and-forget alerts on all their data, so they can be more proactive with their time (and monitor on more nuanced triggers than just what fits on their dashboard page).

Avenue can improve response times to critical problems from several days to real-time by alerting directly on the data sources that customers already use.

data-engineering

saas

developer-tools

Metaplane\ \ Y Combinator LogoW2020\ \ • Acquired • 32 employees • New York City

Metaplane ensures everyone trusts the data that powers your business. Data teams at Bose, Ramp, and Klaviyo use our data observability platform to prevent and detect data issues — before the CEO pings them about weird revenue numbers.

We do this with ML-based anomaly detection, end-to-end column-level lineage, and tools to help prevent incidents before they occur. You can monitor your entire data stack within 30 minutes.

The company is backed by Khosla Ventures, Y Combinator, and the founders of Okta, HubSpot, and Vercel.

data-engineering

developer-tools

saas

Chaos Genius\ \ Y Combinator LogoW2020\ \ • Acquired • 10 employees • Bengaluru

Chaos Genius is a DataOps Observability platform for Snowflake. Enable Snowflake Observability to reduce Snowflake costs and optimize query performance.

cloud-workload-protection

machine-learning

data-engineering

analytics

open-source

Jitsu\ \ Y Combinator LogoS2020\ \ • Acquired • 4 employees • New York City

Jitsu is the fastest, most durable way to collect event data from every source - web, app, email, chatbot, CRM - into your data warehouse. 100% open-source. Purpose built, secure and ready in minutes.

data-engineering

saas

b2b

open-source

Data Mechanics\ \ Y Combinator LogoS2019\ \ • Acquired • 25 employees • Paris, France

Data Mechanics was acquired by NetApp in 2021 and integrated in the Spot.io product portfolio. Our managed Spark-on-Kubernetes platform is live and running under the name Ocean for Apache Spark: https://spot.io/products/ocean-apache-spark/

data-engineering

b2b

saas

open-source

TetraScience\ \ Y Combinator LogoS2015\ \ • Active • 100 employees • Boston

TetraScience provides the world’s first and only R&D Data Cloud, with a mission to transform life sciences R&D, accelerate discovery, and improve human life. Scientists at global pharma and biotech organizations rely on our innovative Tetra Data Platform for easy access to centralized, harmonized, and actionable scientific data to accelerate their digital lab transformation. With best-in-class SaaS performance, a team of industry innovators, and excellent product/market fit, Tetra is positioned to become an iconic life sciences software company.

saas

data-engineering

Yhat (YC W15, pronounced y-hat) was an end-to-end data science platform. Acquired by Alteryx (NYSE:AYX)

data-engineering

machine-learning

enterprise

artificial-intelligence

Scuba\ \ Y Combinator LogoW2013\ \ • Acquired • 51 employees

Scuba is the fast and scalable event-based analytics solution to answer critical business questions about how customers behave and products are used. Interana allows users to analyze and explore the key business metrics that matter most in a data-driven world – such as growth, retention, conversion and engagement – in seconds, rather than the hours or days it often takes with existing solutions. Interana allows customers to discover and investigate these key insights easily through its visual and interactive interface, which makes data analysis a natural extension of everyone’s workflow.

analytics

big-data

data-engineering

data-visualization

BackType\ \ Y Combinator LogoS2008\ \ • Acquired0 • San Francisco

saas

data-engineering

Hottest Startup Categories

Startups by Industry

Startups by Location

Startups Hiring by Location