[go: up one dir, main page]

Reinforcement Learning - now in beta

Post-training, all the way to production

RL and SFT, full-weight and LoRA, on frontier open models. The highest-quality custom models through fast, large-scale experimentation.

Why Together Custom Training

One platform that takes a model from first experiment to production, without leaving the stack.

State-of-the-art quality

Full-weight training at frontier scale, native expert training for MoE architectures, and reinforcement learning validated at the token level. Custom models engineered the way frontier labs build them.

Cost-efficient at scale

Guaranteed throughput on dedicated infrastructure, with on-demand compute when you need it. Concurrent LoRA training and purpose-built stacks keep every GPU working, for industry-leading price-performance.

End-to-end platform

Run concurrent reinforcement learning experiments, validate with high-scale multi-LoRA inference, and merge and quantize for serving at scale. The fastest path from experiment to production, on one platform.

Choose your training method

Both run through a granular Python SDK, with high configurability over every training job.

  • Reinforcement learning

    Optimize your model against a reward.

    GRPO and custom losses
    Reward and KL tracking per run
    Frontier open models, full-weight or LoRA
  • Supervised fine-tuning

    Teach the model from your own demonstrations.

    Your data, your formats
    Instruction and conversation tuning
    Same SDK, same deployment path

Everything you need to reach production

Your code defines the run. Full configurability, dedicated capacity, and one path from training to serving.

    • High-scale experimentation

      Parallel adapters
      One deployment
      Fast test loop

      Run many LoRA adapter experiments in parallel inside one training deployment. Converge on your best model in one fast test loop.

    • Training precision

      Matched computations
      Stable reward curves

      We match the computations between training and inference and minimize any discrepancies, keeping training stable even for the largest runs. Reach the strongest version of your model.

    • Dedicated capacity

      Reserved for you
      Predictable throughput
      Predictable cost

      Your experiments run on capacity that is completely yours. Predictable performance and predictable cost, with your code defining the run.

    • End-to-end platform

      Sandbox evaluation
      Shadow traffic
      Dedicated inference

      Iterate on experiments, deploy multiple versions, and promote the best to production, with a dashboard tracking every run. One stack carries the model straight to serving.

    High-scale experimentation workflow diagram

Backed by frontier research

Every run sits on Together's own training and inference research.

  • Raro
  • Upipe
  • FFT Optimizer
  • Accurary (%) on DeepMath reasoning benchmark

    RARO vs VERIFIER-FREE

    +10% accuracy

    RARO, our Relativistic Adversarial Reasoning Optimization method, learns strong reasoning from expert demonstrations, without verifiers. It outperforms the best verifier-free methods in domains where ground truth doesn't exist, and scales like verifier-based RL.

    Learn more
  • Context parallelism approaches on long-context training

    • Together AI (DCT)
    • Baseline (LD)

    UPipe vs other SOTA Approaches

    82.5% less memory

    Long-context training hits a memory wall at the attention layer. UPipe processes attention heads in smaller chunks, cutting peak activation memory by up to 82.5% — enabling 5M token context lengths on a single 8×H100 node.

    Learn more
  • Context parallelism approaches on long-context training

    • Together AI (DCT)
    • Baseline (LD)

    FFT Optimizer results

    25% less memory

    Fine-tuning large models is memory-hungry. Our FFT-based optimizer replaces expensive SVD projections with fast Fourier transforms, reducing optimizer memory by up to 25% with no loss in training quality.

    learn more

Production-grade
security and data privacy

We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.

Learn More

We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.

  • NVIDIA logo with text Preferred Partner on a black background.
    preferred partner
  • SOC 2 Type II
  • ISO 27001:2022