newsfilter.io
Product Demonstration, Other

GPT-4.5 = Big Model Energy | YC Decoded

  • OpenAI has officially released GPT-4.5, described as the organization's largest and most human-like model to date, representing a significant step in scaling unsupervised learning.
  • The model was previously known internally as "Orion," a project distinct from the "Strawberry" rumors and the O1 reasoning model released in December 2023.
  • Initial market reaction has been muted, as benchmark improvements over GPT-4.0 are viewed as incremental rather than revolutionary.

Performance and Capabilities

  • GPT-4.5 is potentially over 10 times the size of GPT-4.0, leveraging advanced pre-training and post-training scaling.
  • Accuracy on the SimpleQA benchmark increased to 61.9%, a substantial jump from GPT-4.0's 38.4%.
  • Hallucination rates dropped significantly to approximately 37%, down from 61.2% with GPT-4.0.
  • The model excels in emotional intelligence and complex planning, demonstrating a deeper understanding of human intent and nuance.
  • Creative outputs, including email drafting, storytelling, and humor, are rated as more human-like and capable of grasping irony compared to previous iterations.
  • GPT-4.5 outperformed GPT-4.0 and O1 on persuasive power benchmarks titled "Make Me Pay" and "Make Me Say."
  • Evaluation of the model relies heavily on "vibes testing" and human feedback loops due to the subjective nature of assessing writing quality and emotional nuance.

Limitations and Cost

  • Pricing is significantly higher than previous models, costing 30 times more per input token and 15 times more per output token than GPT-4.0.
  • The high cost and architecture make the model unsuitable for large-scale deployment at this time.
  • Performance in structured reasoning domains, such as complex STEM tasks, advanced mathematics, and difficult coding challenges, lags behind specialized reasoning models like O1.

Future Implications and Roadmap

  • OpenAI researchers posit that while unsupervised pre-training scaling continues to yield value, the next major gains will likely come from investing more in inference-time reasoning.
  • Sam Altman predicts the convergence of unsupervised pre-trained models (like GPT-4.5) and specialized reasoning models (like O1) into a unified architecture for GPT-5.
  • The future architecture is expected to combine vast world knowledge, creative fluency, emotional nuance, and advanced reasoning within a single system.

Announcements

  • YC is hosting its first AI Startup School in San Francisco on June 16–17, featuring confirmed speakers including Elon Musk, Satya Nadella, Sam Altman, and Andrej Karpathy.
  • The conference is free for computer science graduate students, undergraduates, and new graduates in AI, with travel to SF covered for applicants.