AI 行业热点
🎙️ 播客精选
Why Traditional Benchmarks Fail Modern AI Models with OpenAI Research Scientist Noam Brown
No Priors · 2026-06-26
Speaker 1 | 00:00 - 00:13
With GBT three, you couldn’t scale test time compute. Like, if you gave it a budget of $10,000,000 and said, okay. Well, let’s see what GBT three can do. It really can’t do that much. The precarious frameworks and responsible scaling policies, they don’t really account for the amount of test time compute.
Speaker 1 | 00:13 - 00:32
They just say, okay. Well, what’s the capability of the model? The problem is we’re in a world now where the capability of the model is a function of how much money you put into it, basically. If you give it a budget of $10,000, it can do a…
🐦 X/Twitter 热点
Swyx (@swyx)
- minor milestone in the growth of swyx inc:
we took over my new media lab today
it will be the new home for engineer-creatives in san francisco; a third place to make; a finishing school for technical storytellers; a place to inspire the inspirational.
to our complete surprise; it came with a datacenter rack randomly set up and wired up! need advice on what to put in this thing. @alexocheema halp [72 ❤️ 2 🔄]
- we have been scaling without slop by working with aligned domain experts to add coverage
with both oai and ant launching multi-billion dollar services arms, it’s clear that FDE is one of the most in demand disciplines on earth, but I have never done the job
it’s been an absolute pleasure working with Basil on our first ever AI FDE miniconference!
see at next week [46 ❤️ 2 🔄]
Thibault Sottiaux (@thsottiaux)
- We are giving all Codex users a usage reset on the house. Should be showing in your accounts in the next few hours.
We have applied some mitigations, but our investigation hasn’t shown users being impacted at large. We are continuing to monitor the situation. [4072 ❤️ 208 🔄]
- [255 ❤️ 2 🔄]
- [394 ❤️ 10 🔄]
Peter Yang (@petergyang)
- So let me get this straight:
- We publish frontier models
- They get distilled into cheap open source models
- US companies adopt the same open source models because they’re good enough and much cheaper
- We start gating access to frontier models
What next? US companies innovate less? Open source models become more attractive? [597 ❤️ 19 🔄]
- From what I’m seeing, alot of the money has moved to services (with some software bundled), not software.
People want outcomes, not tools.
It’s feels really hard to build a pure-play software company that’s more valuable to people or companies than just using Codex/Claude Code with a bunch of personal skills and agents.
Thoughts? [115 ❤️ 9 🔄]
- Small things I wish Claude Code had:
Bring back ability to steer conversations while Claude is working
Make mobile remote control for all threads on by default
The shortcut keys seem only accessible if you have the sub-menu open? Consider supporting “cmd + key” so we can hotkey to different threads with the keyboard. If I hold down cmd I should see all the shortcuts in the UI.
Let me drag and drop to re-arrange my projects on left nav [92 ❤️ 4 🔄]
Nan Yu (@thenanyu)
- Patient zero [1 ❤️ 1 🔄]
- Socks don’t have to match. It’s cuter that way anyway [1 ❤️]
- Secret level 6: there’s a problem, but it’s not worth solving, so leave it alone.
This is why org with a ton of level 1s can win. They don’t get distracted by side quests [129 ❤️ 4 🔄]
Cat Wu (@_catwu)
- split screen is one of my fave claude code on desktop features! [268 ❤️ 10 🔄]
Guillermo Rauch (@rauchg)
- Agents are particularly hard-to-debug software.
For one, and by design, AI models behave in non-deterministic ways. Even two identical prompts don’t always yield the same output.
But agents are also complex distributed systems. They involve multiple steps of computation across functions and sandboxes, touching dozens of API services that can go down, rate limit you, etc.
Nailing down observability out of the box for on Vercel was a key priority for the team, and the feedback so far has been 🔥↓ [270 ❤️ 8 🔄]
- I made a @hyperframes_ video of how I found out this morning [68 ❤️ 4 🔄]
- The UI for AI is here. It’s @shadcn [1738 ❤️ 38 🔄]
Aaron Levie (@levie)
- Step one complete [140 ❤️ 2 🔄]
- GPT-5.6 is real and looks very strong. Going to be very strong for knowledge worker tasks that require heavy tool use and long running agents doing work. We’re not hitting any walls in AI progress right now. [239 ❤️ 18 🔄]
Garry Tan (@garrytan)
- This is honestly no way to release a model and continued development and release this way is a solid way to salt the ground and kill all innovation by small startups [586 ❤️ 56 🔄]
- Don’t be a mid startup
It’s true [1378 ❤️ 75 🔄]
Matt Turck (@mattturck)
- World Cup scoring strategies:
Argentina: pass the ball to Messi
Portugal: pass the ball to Ronaldo
England: pass the ball to Kane
Spain: pass the ball to Lamal
France: pass the ball to Mbappé or Dembélé or Olise or Doué or Barcola or Cherki or Thuram or [22903 ❤️ 1706 🔄]
- For once, this tweet aged well 😯 [4 ❤️]
- If Dembélé is now fully confident, on top of everything else the team has going on, good luck to all the other nations [52 ❤️]
Zara Zhang (@zarazhangrui)
- [88 ❤️ 7 🔄]
- “You do not need God to write your emails” [93 ❤️]
- Btw I used Borumi to create this video, and it’s genuinely the most underrated video recording/editing tool out there
It’s like Screen Studio + Descript + CapCut all in one
It’s so underrated; they don’t even have an X account I can tag [212 ❤️ 11 🔄]
Nikunj Kothari (@nikunj)
- Now that the dust is settled, what’s really funny to me about the taste commentary is it’s coming from a LOT of people who never have built anything..
You can’t achieve (or refine) taste without being in the arena.
It’s like a chef cooking and expecting their first dish to be good on the first go.
But, when you give it a 100 iterations and maybe cook a 100 dishes, then the 101st one has imbued all the lessons from those 10,000 iterations.
But, if you over fit without much variation, then people get tired of it since well they want something new as well.
The best people are able to learn, imbue and innovate by breaking patterns consistently.
It’ll be interesting to see how AI does the same - and unlike most people I think it has a fair shot at actually building taste.
End rant. [40 ❤️ 3 🔄]
- Written for the seed founder who’s building something really important in a not-so-hot category..
I see you 🙏 [137 ❤️ 4 🔄]
Peter Steinberger (@steipete)
- I love how Apple notarization breaks multiple times a year until I manually log in and accept some new legal agreements. [947 ❤️ 8 🔄]
Dan Shipper (@danshipper)
- full blog post here: [9 ❤️]
- BREAKING: OpenAI announced GPT-5.6 Sol!
As of today, by U.S. government directive, access is limited to only ~20 pre-approved companies and @every is not on the list. This appears to be a temporary situation while the government races to figure out a long-term policy for releasing frontier models with advanced capabilities.
I understand and applaud the need for some government oversight in making American infrastructure resilient to cyberattacks and other potential new threats from the misuse of these models. However, I also strongly believe that widespread democratic access to frontier models is absolutely necessary to our country’s leading position in the AI race. It’s also critical for allowing American workers to keep pace with the new skills they need to be productive in this era.
A world where advanced models are locked up only for use by the employees of AI giants and a select few companies is one where ambitious students, independent builders, and working professionals are denied the tools they need to learn, create, and compete to their fullest potential.
Also, speaking for @every and the community of developers and writers who share early access with us: Our job is to test these tools early so that we can prepare Americans for how to use them to their fullest extent. If we lose access, we lose the ability to do that important job well.
Thankfully, both OpenAI and the government seem to be working to make broad access available soon.
We’ll be ready to vibe check when that happens :) [231 ❤️ 13 🔄]
Aditya Agarwal (@adityaag)
- An interesting side effect of AI is that I have zero tolerance for shallow interactions with humans.
I really deep connection and depth in my relationships. More ways to do that are so valuable.
And for everything else: give me agents.
I suspect our world will become both smaller and more rich in it’s relationship depth. [67 ❤️ 5 🔄]
Sam Altman (@sama)
- not quite all-you-can-eat tokens, but we are working on it [753 ❤️ 23 🔄]
- team cooked, spicily [2753 ❤️ 76 🔄]
- in other news, we updated the 5.5 instant model used in chatgpt this week.
i like its vibes. [4164 ❤️ 79 🔄]
由 Follow Builders 自动生成 · 2026-06-27