AI 行业热点
🐦 X/Twitter 热点
Swyx (@swyx)
- the reception is unlike anything i thought possible for a 2026 OAI launch [1 ❤️]
- [18 ❤️ 2 🔄]
- sorry for the radio silence folks - got sucked into extreme LLM psychosis. but i can confidently say we have crossed over into a new age of AI Engineering and we are never, ever, looking back.
this isnt even EVERYTHING i did with Astra but i’ll append more reports as I publish them on LS! [780 ❤️ 35 🔄]
Boris Cherny (@bcherny)
- Your input needed: would you use this?
This is an early look at how we’re thinking about making Claude Code way more extensible. It’s a little crazy, and very exciting.
More details here: [880 ❤️ 49 🔄]
Thibault Sottiaux (@thsottiaux)
- We will give one banked reset for every day you don’t have access to Astra on your paid ChatGPT plan, starting today. Team is moving mountains to give access as fast as we can.
First one will land in ~ 3 hours. There is still time to create your account if you don’t have one. [38250 ❤️ 2955 🔄]
- We are going to need a different AGI benchmark. Where is the goalpost moving next? [11392 ❤️ 560 🔄]
- Across the plans Astra will be included in the normal usage allocation and you will be able to use 100% of it towards Astra. [7943 ❤️ 273 🔄]
Peter Yang (@petergyang)
- I live in Codex and think it’s the best software shipped in the past 5 years.
But the combination of influencers posting non-stop about how great Astra is and paid users unable to get access at the same time is pretty rough IMO.
I realize I’m an “influencer” now too and maybe am just jealous. I also recognize there are multiple parties involved in these decisions.
But just wanted to be real. Back to working with good ol’ Sol and what’s left of my Fable limits. [599 ❤️ 9 🔄]
- Wait so we don’t get to play with Astra today? [528 ❤️ 4 🔄]
- I have one codex thread that’s still working and i’m hanging onto it for dear life lol [75 ❤️]
Nan Yu (@thenanyu)
- Time to touch grass [24 ❤️]
Madhu Guru (@realmadhuguru)
- I was just in a meeting where someone casually used
“load-bearing argument”,
“that’s the spine of our plan”,
“one honest callout “
in conversation.
I think the machines have successfully RL’d us [59 ❤️ 3 🔄]
- The hard part with becoming more ambitious is changing the how.
We’re living in a time where AI and market conditions make asymmetric outcomes far more possible.
That applies to scale, velocity, breadth of product, personal career and wealth goals.
You need to drop the ideas that aren’t serving you - rethink team structure, roadmapping.
On the personal front, rethink how you see yourself and how you engage with people.
The blocker is often inertia - institutional, personal.
“We’ve always done it this way.”
“That feels too risky.”
“That would be cringe.”
Write down current goals for your product or for you.
Then ask what 100X would look like.
Challenge yourself on the habits and assumptions you need to drop to get there. [111 ❤️ 4 🔄]
Amanda Askell (@AmandaAskell)
- I still don’t know how to convey my enthusiasm in a way that Americans will read as enthusiasm and not something more like unwilling resignation. Has any British person achieved this? If so, how is it done? [427 ❤️ 14 🔄]
Thariq (@trq212)
- We’re working on making Claude Code way more hackable, give us feedback! [329 ❤️ 8 🔄]
Amjad Masad (@amasad)
- Marvin Minsky wrote about this in The Emotion Machine. Emotions are a core part of human intelligence, not some epiphenomenal side effect of human evolution (which presumably many AI folks believe it is).
Minsky describes a selector of sorts for different thinking strategies. [61 ❤️ 4 🔄]
- GPT-6 is a major jump in capabilities and will unlock new use-cases. Will launch on Replit very soon for you to try it! [1124 ❤️ 40 🔄]
- Really cool to see Foster City, where Replit is HQ’d, councilwoman is building on Replit to help the community! [77 ❤️ 4 🔄]
Guillermo Rauch (@rauchg)
- “Feedback is a gift” has always been a favorite mantra of mine and a strongly-held belief.
Now it’s fact. Each piece of feedback is someone gifting you a prompt for you to give your agents to improve your product ☺️
Thankful for everyone who takes time (or tokens) to critique our products, write to us or forward your agents’ transcripts. [257 ❤️ 17 🔄]
- If you intern at Vercel, you can work on incredible projects like this one. Improving Next.js chunking can yield massive efficiency improvements at internet scale. Excellent writeup by @sam_poder [430 ❤️ 14 🔄]
- Run
▲ ~/ vercel ai-gateway coding-agents setup
Points all coding agents to AI Gateway, getting you 100% uptime, observability, budgets, and ease of switching. [171 ❤️ 6 🔄]
Aaron Levie (@levie)
- GPT-6 Astra is out. We’ve been testing the model in early preview on our enterprise complex work eval at Box. It is now the best model we’ve ever tested on our expanded and hardest test set.
Overall, GPT-6 Astra offers a breakthrough level of capability in coding, analytics, logic, and domain specific knowledge for dealing with complex enterprise knowledge work. In our eval, we use the GPT-6 Astra model in the Box Agent and give it a set of difficult tasks that represent real-world work in various industries. These tasks often will take people hours to execute properly and require deep domain expertise.
GPT-6 Astra scored 77% overall vs. 74% with GPT-5.6 Sol, representing frontier performance. But the story is in some of the bigger gains on individual task types that are incredibly complex across. Here are a few examples across the tests that show how much of an impact this will be in different industries:
Media and entertainment (48% → 100%, +52 points). Ranking film genres and countries by profitability across a year of production data from four teams, where the trick is to apply the following year’s tax-incentive corrections without double-counting them. GPT-6 Astra was perfect on every attempt. Sol had the rankings right but the underlying ratios were wrong.
Technology (69% → 97%, +28 points). Choosing which region to fund first, from a stack of performance and infrastructure documents that never state the metric the brief asks for. GPT-6 Astra flagged the gap, labelled its own figure a proxy, and caught a growth claim that didn’t match the numbers beneath it. GPT-5.6 Sol reported the proxy as the real thing.
Legal (69% → 93%, +24 points). Reviewing an NDA against a company’s own contracting policy, where reaching the right verdict isn’t enough; the answer has to cite the provision behind it. Both models declined to approve the draft, but only GPT-6 Astra separated whether the liability cap’s structure was permissible from whether its amount was defensible, and pointed to the policy language that settles it. GPT-5.6 Sol argued the same conclusion without citing the provision, which is a critical error in the legal industry.
Healthcare (53% → 77%, +23 points). Auditing a batch of radiology reports for errors and rating how serious each one is. GPT-6 Astra caught a terminology error in a knee-imaging report that GPT-5.6 Sol missed on both attempts, and was more reliable at grading severity rather than just flagging that something was wrong.
Energy (82% → 97%, +15 points). Building a consumption report on two facilities from a year of meter logs, where only one of the two sites is actually missing data. GPT-6 Astra found it and drew a line most reviewers wouldn’t: the other site’s odd solar readings are a measurement problem to investigate, not a gap in the record. GPT-5.6 Sol reported both sites as incomplete.
Astra clearly is going to offer a meaningful jump in powering and orchestrating enterprise workflows. We’ll be making GPT-6 Astra as an option in the Box AI Studio for customers to build agents with shortly as it continues to roll out. [901 ❤️ 82 🔄]
- Another huge moment for open weights AI. The infra platforms are working. The models are getting better. The ecosystems are being deeply invested in. The business models are working. [116 ❤️ 9 🔄]
Garry Tan (@garrytan)
- Grok images quite impressive tbh [12 ❤️]
- People who never had proper jobs or ever ran businesses can’t figure out that markets exist
“Prices, how do they work?” [116 ❤️ 3 🔄]
- #NewProfilePic
Photo by Oliver Covrett [431 ❤️ 3 🔄]
Matt Turck (@mattturck)
- We talked about it on the pod here:
- This is wild. ARC-AGI was built to resist the LLM scaling paradigm
o1 despite early reasoning struggled mightily in 2024 with 18%
Then ARC-AGI-3, an even harder test, launched in 2026. Frontier AI was at 0.5%
And now Astra just completely saturated it (w/ its native harness) [66 ❤️ 2 🔄]
Zara Zhang (@zarazhangrui)
- Grok Bot is what OpenClaw should have been [84 ❤️]
- I wish more founders made more raw screen recordings of real product interfaces and the thinking behind them, rather than polished high-production launch videos [779 ❤️ 48 🔄]
Nikunj Kothari (@nikunj)
- A little behind the scenes of how I made this - it was mostly autonomous, with me spending <20 minutes of active time:
On my drive this morning, I voice memo-ed to Claude this idea and asked it to generate a spec (image 1). I made some visual changes, but the spec was pretty good in one shot.
I gave that spec Codex /goal with 5.6 Sol High and two API keys (reactor and gemini - and yes I rotated them)
It mostly ran autonomously for a few hours and created a pretty good first draft (image 2)
I gave it detailed scene by scene feedback with another /goal and it was able to get it to the place that you see above (image 3)
Total cost:
- Reactor: ~$17
- Nano Banana: $4
- Codex: on the sub, and used only 14% of weekly usage [4 ❤️]
- The Collective, a short film.
I still think a lot of people don’t understand the gravity and the details about the @OpenAI x @huggingface incident..
Primarily, because it’s so complex to understand the technical reasons, and partly because it reads like a science fiction book.
So I used @claudeai Fable 5.1 to reason through the latest @dwarkesh_sp episode with @ajeya_cotra and give me a scene by scene narrative of what a short film could look like.
I then used @ChatGPT Codex + MiniMax Fast H3 on @reactorworld + @NanoBanana images to come up with this film.
I was relatively surprised on how good this turned out to be in mostly one-shot (even without Astra), and hope you can share it with folks who need a visual explainer.
Disclaimer: I tried to fact check this as much as possible using a variety of sources, but there’s still a lot of open questions here that might trigger some folks. I’m going by what’s reported, so please flag any errors! [78 ❤️ 6 🔄]
- None of these chief of staff products work for busy people because they are missing more than half the knowledge that is currently stored in a walled garden (aka your phone)..
Only when you are able to combine it, get folks to train it to teach what’s important, give it episodic memory and push actual proactive things - you can claim to be a real chief of staff.
Till then you’re just a sparkling GSuite & Slack wrapper. Useful, will save you time but definitely not a chief of staff 🙈 [85 ❤️ 3 🔄]
Peter Steinberger (@steipete)
- See you there! Will talk about how we build in the open and multiplayer agents. [65 ❤️ 3 🔄]
- Having a claw in your group chat is so useful! [156 ❤️ 6 🔄]
- brilliant fit. [169 ❤️ 2 🔄]
Dan Shipper (@danshipper)
- how you know you did a good vibe check [39 ❤️]
- VIBE CHECK: GPT-6 ASTRA [26 ❤️ 2 🔄]
- @every read our full vibe check: [10 ❤️]
Aditya Agarwal (@adityaag)
- We have quite the lineup of speakers this Fall.
Waymo, Physical Intelligence, Anduril and Applied Intuition.
Let’s go. [51 ❤️ 2 🔄]
- The single biggest issue with using agents today is speed.
Imagine if the speed were 10-100x faster – the interaction pattern and depth of usage would be vastly different. [66 ❤️]
Sam Altman (@sama)
- We are also excited! [2511 ❤️ 60 🔄]
- first, sorry for the messy rollout.
second, when we screw up, we try to make it right.
third, we should be able to begin broad rollout to API customers and chatgpt subscribers in the near future. as usual we will start with pro subscribers. [12387 ❤️ 431 🔄]
- Also, this is my favorite OpenAI video so far. It makes me excited for the future! [15158 ❤️ 724 🔄]
由 Follow Builders 自动生成 · 2026-09-04