Skip to main content
KYM

Detected Skills:

Reinforcement Learning (100% match)

Found 0 registries and 15 entities for "Reinforcement Learning"

All Agents & Models

OpenAI: o3 Pro

model

The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more compute to think harder and provide consistently...

Reinforcement Learning Code Execution Text Generation Image-Text-to-Text
openai Score: 0

OpenAI: o1-pro

model

The o1 series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o1-pro model uses more compute to think harder and provide...

Reinforcement Learning Code Execution Text Generation Image-Text-to-Text
openai Score: 0

OpenAI: o1

model

The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with large-scale reinforcement learning to reason...

Text Generation Question Answering Text-to-Image Reinforcement Learning +3 more
openai Score: 0

deepseek-ai/deepseek-r1

model

A reasoning model trained with reinforcement learning, on par with OpenAI o1

Text Generation
deepseek-ai Score: 0

Qwen: Qwen3 Max Thinking

model

Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it...

Text Generation
qwen Score: 0

DeepSeek R1 (Fast)

model

DeepSeek R1 (Fast) is the speed-optimized serverless deployment of DeepSeek-R1. Compared to the DeepSeek R1 (Basic) endpoint, R1 (Fast) provides faster speeds with higher per-token prices, see https://fireworks.ai/pricing for details. Identical models are served on the two endpoints, so there are no quality or quantization differences. DeepSeek-R1 is a state-of-the-art large language model optimized with reinforcement learning and cold-start data for exceptional reasoning, math, and code performance. The model is identical to the one uploaded by DeepSeek on HuggingFace. Note that fine-tuning for this model is only available through contacting fireworks at https://fireworks.ai/company/contact-us.

Text Generation
fireworks Score: 0

DeepSeek R1 (Basic)

model

DeepSeek R1 (Basic) is the cost-optimized serverless deployment of DeepSeek-R1. Compared to the DeepSeek R1 (Fast) endpoint, R1 (Basic) provides lower per-token prices with slower speeds, see https://fireworks.ai/pricing for details. Identical models are served on the two endpoints, so there are no quality or quantization differences. DeepSeek-R1 is a state-of-the-art large language model optimized with reinforcement learning and cold-start data for exceptional reasoning, math, and code performance. The model is identical to the one uploaded by DeepSeek on HuggingFace. Note that fine-tuning for this model is only available through contacting fireworks at https://fireworks.ai/company/contact-us.

Text Generation
fireworks Score: 0

KAT Dev 32B

model

KAT-Dev-32B is an open-source 32B-parameter model for software engineering tasks. It is optimized via several stages of training, including a mid-training stage, supervised fine-tuning (SFT) & reinforcement fine-tuning (RFT) stage and an large-scale agentic reinforcement learning (RL) stage.

fireworks Score: 0

Llama 3 8B

model

Llama 3 is an auto-regressive language model that uses an optimized transformer architecture. The tuned versions use supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to align with human preferences for helpfulness and safety.

Text Generation
fireworks Score: 0

MiniMax-M2.5

model

MiniMax M2.5 is built for state-of-the-art coding, agentic tool use, search, and office work, extensively trained with reinforcement learning across hundreds of thousands of real-world environments to plan like an architect and generalize across unfamiliar scaffolding and tools. It delivers significantly faster task completion, improved token efficiency, and exceptional cost-effectiveness, making it well-suited for production-scale agentic applications and complex, multi-step workflows.

fireworks Score: 0

OpenChat 3.5 0106

model

OpenChat is an innovative library of open-source language models, fine-tuned with C-RLFT - a strategy inspired by offline reinforcement learning.

Text Generation
fireworks Score: 0

DeepSeek V3.2

model

We introduce DeepSeek-V3.2, a next-generation foundation model designed to unify high computational efficiency with state-of-the-art reasoning and agentic performance. DeepSeek-V3.2 is built upon three core technical breakthroughs: • DeepSeek Sparse Attention (DSA): A new highly efficient attention mechanism that significantly reduces computational overhead while preserving model quality, purpose-built for long-context reasoning and high-throughput workloads. • Scalable Reinforcement Learning Framework: DeepSeek-V3.2 leverages a robust RL training protocol and expanded post-training compute to reach GPT-5-level performance. Its high-compute variant, DeepSeek-V3.2-Speciale, surpasses GPT-5 and demonstrates reasoning capabilities comparable to Gemini-3.0-Pro. • Large-Scale Agentic Task Synthesis Pipeline: To enable reliable tool-use and multi-step decision-making, we develop a novel agentic data synthesis pipeline that generates high-quality interactive reasoning tasks at scale, greatly enhancing the model’s

Text Generation
deepseek Score: 0

MiniMax M1

model

MiniMax-M1: The World's First Open-Weight, Large-Scale Hybrid Attention Inference Model MiniMax-M1 adopts a Mixture of Experts (MoE) architecture and integrates the Flash Attention mechanism. The model contains a total of 456 billion parameters, with 45.9 billion parameters activated per token. Natively, the M1 model supports a context length of 1 million tokens—8 times that of DeepSeek R1. Additionally, by combining the CISPO algorithm with an efficient hybrid attention design for reinforcement learning training, MiniMax-M1 achieves industry-leading performance in long-context reasoning and real-world software engineering scenarios.

Text Generation
minimaxai Score: 0

Baichuan M2 32B

model

Baichuan-M2 is a medically-enhanced reasoning model specifically designed for real-world medical reasoning tasks. We begin with real-world medical questions and conduct reinforcement learning training based on a large-scale verifier system. While maintaining the model's general capabilities, the medical effectiveness of Baichuan-M2 has achieved breakthrough improvements. Baichuan-M2 is currently the world's best open-source medical model. On the HealthBench Benchmark, it surpasses all open-source models, including GPT-OSS-120B, as well as many cutting-edge closed-source models. It is the open-source model closest to GPT-5 in terms of medical capabilities. Our research demonstrates that a robust verifier is crucial for aligning model capabilities with real-world applications, and an end-to-end reinforcement learning approach fundamentally enhances the model's medical reasoning abilities. The release of Baichuan-M2 represents a significant advancement in the field of medical artificial intelligence, pushing t

Text Generation
baichuan Score: 0

GLM-4-32B-0414

model

GLM-4-32B-0414 is the latest open-source model in the GLM series, featuring 32 billion parameters. Its performance is comparable to OpenAI's GPT series and DeepSeek's V3/R1 series, while also supporting highly user-friendly local deployment capabilities. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including a large amount of reasoning-type synthetic data, which laid a solid foundation for subsequent reinforcement learning extensions. In the post-training stage, in addition to human preference alignment for dialogue scenarios, the research team enhanced the model’s performance in instruction following, engineering code, and function calling using techniques such as rejection sampling and reinforcement learning, thereby strengthening the atomic capabilities required for agent tasks. GLM-4-32B-0414 has achieved strong results in engineering code generation, artifact creation, function calling, search-based question answering, and report generation. On several benchmarks, its performance appr

Text Generation
thudm Score: 0