Product Demonstration, Other
GPT-4.5 = Big Model Energy | YC Decoded
- OpenAI has officially released GPT-4.5, described as the organization's largest and most human-like model to date, representing a significant step in scaling unsupervised learning.
- The model was previously known internally as "Orion," a project distinct from the "Strawberry" rumors and the O1 reasoning model released in December 2023.
- Initial market reaction has been muted, as benchmark improvements over GPT-4.0 are viewed as incremental rather than revolutionary.
Performance and Capabilities
- GPT-4.5 is potentially over 10 times the size of GPT-4.0, leveraging advanced pre-training and post-training scaling.
- Accuracy on the SimpleQA benchmark increased to 61.9%, a substantial jump from GPT-4.0's 38.4%.
- Hallucination rates dropped significantly to approximately 37%, down from 61.2% with GPT-4.0.
- The model excels in emotional intelligence and complex planning, demonstrating a deeper understanding of human intent and nuance.
- Creative outputs, including email drafting, storytelling, and humor, are rated as more human-like and capable of grasping irony compared to previous iterations.
- GPT-4.5 outperformed GPT-4.0 and O1 on persuasive power benchmarks titled "Make Me Pay" and "Make Me Say."
- Evaluation of the model relies heavily on "vibes testing" and human feedback loops due to the subjective nature of assessing writing quality and emotional nuance.
Limitations and Cost
- Pricing is significantly higher than previous models, costing 30 times more per input token and 15 times more per output token than GPT-4.0.
- The high cost and architecture make the model unsuitable for large-scale deployment at this time.
- Performance in structured reasoning domains, such as complex STEM tasks, advanced mathematics, and difficult coding challenges, lags behind specialized reasoning models like O1.
Future Implications and Roadmap
- OpenAI researchers posit that while unsupervised pre-training scaling continues to yield value, the next major gains will likely come from investing more in inference-time reasoning.
- Sam Altman predicts the convergence of unsupervised pre-trained models (like GPT-4.5) and specialized reasoning models (like O1) into a unified architecture for GPT-5.
- The future architecture is expected to combine vast world knowledge, creative fluency, emotional nuance, and advanced reasoning within a single system.
Announcements
- YC is hosting its first AI Startup School in San Francisco on June 16–17, featuring confirmed speakers including Elon Musk, Satya Nadella, Sam Altman, and Andrej Karpathy.
- The conference is free for computer science graduate students, undergraduates, and new graduates in AI, with travel to SF covered for applicants.