
Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community
Quick Answer
Together AI and Y Combinator have launched the first dedicated GPU cluster for YC's AI-native startups, addressing the compute bottleneck that many face.
Quick Take
Startups can now access flexible, cost-effective GPU resources without long-term commitments, enabling rapid scaling and innovation.
Key Points
- The GPU cluster provides dedicated access for YC's AI-native startups.
- Startups can provision GPUs for short-term needs with long-term rates.
- Together AI supports over 8,000 customers, enhancing AI application development.
- Founders manage GPU usage directly through Together's self-service portal.
- The cluster is fully utilized, allowing companies to plan compute needs in advance.
DeepSignal Analysis
What happened
Together AI and Y Combinator have launched a dedicated GPU cluster aimed at supporting YC's AI-native startups. This initiative addresses the compute bottleneck that many startups face, allowing them to access GPU resources flexibly and cost-effectively without long-term commitments.
Key evidence
- Together AI has built a dedicated GPU cluster specifically for YC portfolio companies, enabling them to access compute resources for AI applications.
- Startups can reserve and manage their own GPU resources through Together's self-service portal, allowing for rapid scaling without routing through YC.
- The cluster is currently running at full utilization, indicating strong demand from YC startups for dedicated compute resources.
Why it matters
Access to compute resources is critical for AI startups, particularly as model complexity and costs rise. This partnership aims to alleviate financial pressures on startups by providing flexible GPU access, which could enhance innovation and competitiveness in the AI sector. By lowering barriers to entry for compute resources, Together AI and YC are positioning themselves as key enablers for emerging AI companies.
📖 Reader Mode
~3 min readToday, Together AI and Y Combinator (YC) are announcing a partnership to deliver the first dedicated YC GPU cluster, giving YC's portfolio of AI-native startups easier access to the compute they need to build and scale.
Compute has become the biggest bottleneck
Breakthroughs in AI have led to a new generation of companies building AI applications, and a key input to building any of these apps is access to compute. More and more startups struggle to get this access in a timely, cost-effective way.
It used to be simple. A startup could spin up GPU instances on demand and scale from there. Today, as model quality and token value keep rising, just securing capacity is one of the hardest problems a young company faces, let alone getting good pricing on it.
For many startups, the upfront cost to secure two years of compute exceeds their entire cash balance, forcing a choice: Raise a round just to fund a compute contract, or go without the capacity to compete.
Flexible, dedicated access to compute
Together AI and Y Combinator have partnered to solve that specific problem. Together AI has built a dedicated cluster and developer experience just for YC Portfolio companies to quickly and cost-effectively get access to compute for their AI needs across inference and training. Instead of long-term commitments, startups can spin up GPUs for short-term sprints while benefiting from long-term rates.
The cluster supports the full range of needs across YC's portfolio, from teams requiring single-node compute to companies scaling up as their needs grow. Additionally, YC start up founders benefit from the learnings and best practices that Together researchers, engineers, and customer experience teams bring from working with leading AI-native companies.
Startups reserve, provision, and manage their own GPUs directly through Together’s self-service portal, with their own billing support, so scaling compute never has to route through YC. Founders get GPUs ready in minutes, and stay in control of their own usage from day one.
The cluster is running at full utilization today, while individual companies can still plan their compute needs months in advance.
A natural fit
Together AI was founded four years ago on the belief that generative AI would become foundational, and built a cloud service spanning the full generative AI lifecycle. It now works with over 8,000 customers, including Cursor, Decagon, Cognition, and Eleven Labs.
As models improve, the value of every token produced keeps climbing, and so does the cost of compute behind it. Together's research, including work on attention mechanisms and the Mamba architecture now used in models like NVIDIA’s Nemotron, is aimed squarely at that problem: faster inference and better unit economics for every workload on the cluster, from early-stage teams to companies already operating at scale.
YC has long been one of the largest seed funders of research-driven companies, and securing compute has become essential to attracting and supporting the best founders.
Both organizations share a belief that founders do their best work with access to the same resources a much bigger company would have. Compute is simply the latest one.
What's next
Together AI and YC plan to expand the cluster, and Together will keep supporting companies as they graduate from the batch and scale with whatever setup makes sense long term. As those needs grow, founders can also tap into the rest of Together’s platform, from inference to fine-tuning and training, all on the same infrastructure they already know.
If you're building an AI-native startup and want to learn more, contact us.
YC's next batch is currently taking applications for their Fall 2026 cycle. YC is especially excited about funding companies developing new research breakthroughs that require GPU compute. Apply at ycombinator.com/apply.
— Originally published at together.ai
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Together AI
See more →
Open, convenient and predictable: Introducing Provisioned Throughput
Together AI introduces Provisioned Throughput, offering guaranteed inference capacity for MiniMax M3 and GLM-5.2 at $0.05 per PTU per minute, achieving costs up to 90% lower than Claude Opus 4.8. This new model provides predictable pricing and a 99% uptime SLA, catering to companies transitioning to open weight models for production workloads.

