AI 行业热点

🎙️ 播客精选

Parallel’s Parag Agrawal: Building a New Web for AI Agents

Training Data · 2026-08-25

Speaker 1 | 00:00 - 00:40
Our view at Parallel is that human click data is a bug, and agent doing work with search should rely on agent feedback, not human feedback. We believe that these models are really good at compressing information and we can benefit from a lot of the research that have gone into building models and apply it to search indexing and ranking. And so you can now make many, many arguments. And that’s the arguments we made back then that actually now it’s way more tractable as a problem because of the existence of agents, not just as in technology, but as a distinct customer.

🎧 收听完整节目


🐦 X/Twitter 热点

Swyx (@swyx)

  • PSA: do not use codex “locked use” capabilities right now. it is currently relying on unstable mac features and has completely locked me out of my macos keychain twice this week.

thx @_chenglou for linking to apple developer forums acknowledging this is a “known bug”. just avoid. ofc, would be nice to do everything in cloud, but cloud isn’t there yet. [5 ❤️ 1 🔄]

Boris Cherny (@bcherny)

  • A small improvement: memory is now simpler and more powerful. Enjoy! [1429 ❤️ 46 🔄]

Thibault Sottiaux (@thsottiaux)

  • Good products take time. At least 34 days. [2034 ❤️ 62 🔄]
  • I want to see the polymarket [367 ❤️ 4 🔄]
  • So much demand for this one. Works similar to the Pro $100 plan but designed for teams and small companies.

✅ All ChatGPT, ChatGPT Work, and Codex features
✅ Connect to Google Workspace, Slack, GitHub, Microsoft 365, and more
✅ Secure workspace with SAML, SSO, and MFA
✅ Centralized billing and administration
✅ Usage analytics and spend controls
✅ No 5h limits [2926 ❤️ 110 🔄]

Peter Yang (@petergyang)

  • Here’s an example brief that the /fuck-cancer skill creates and updates.

I’ve found it really useful to have a single source of truth doc with patient information, what to do next, what we know, definitions of medical terms, and an update log.

📌 Get the free skill here: [101 ❤️ 11 🔄]

  • Very excited to attend @OpenAI Dev Day.

ChatGPT / Codex is basically the OS I use for everything I do on my computer.

Can’t wait to see what updates the team has in store. [117 ❤️ 7 🔄]

  • Today, I’m open-sourcing /fuck-cancer, an AI skill that helps patients and caregivers navigate cancer diagnosis and treatment and advocate for themselves and their loved ones.

Here’s what patients and caregivers have told me since I shared my mom’s story:

“You have to be a huge patient advocate. Ask questions and push for answers, biopsies, and proper testing.”

“I need help navigating difficult conversations and advocating for myself.”

“The volume of doctors, documents, and insurance paperwork can quickly become overwhelming.”

I built the skill to help with these problems. It creates and updates a practical brief with five sections:

  1. Patient and care-team information for easy reference during calls

  2. What to do next, limited to three specific actions

  3. What we know, separating confirmed facts from what remains unclear

  4. Medical terms explained in plain English

  5. A care log with recent updates and decisions

It builds this brief from documents, and context you provide. When research is needed, it uses trusted sources such as the National Cancer Institute and the ClinicalTrials gov API.

I use it with ChatGPT/Codex and Claude Code to prepare for conversations and research. It can save the brief locally as a Markdown file or update a shareable Google Doc so the whole family can work from the same information.

📌 Get the free, open-source skill here:

If you find it useful, please ⭐ the repo and share it with someone going through this BS disease. [1389 ❤️ 158 🔄]

Madhu Guru (@realmadhuguru)

  • Full eval series so far:

1/ Getting started with evals

2/ Quality first, then cost

3/ Failure modes taxonomy

4/ Laddered eval strategy

5/ Tyranny of the average

6/ Hill climbing on evals

7/ The Goldilocks principle

8/ Discriminatory power of evals

9/ The Eval Roadmap Problem

  • How to build great evals — Part 9

The Eval Roadmap Problem

Most evals fail because teams treat them as static artifacts while their users expectations and behaviors have evolved.

Your evals need a roadmap that evolves with your product and actual usage patterns.

Take a financial research agent. Here is how use cases will evolve over time.

Early user: Summarize this 5-page earnings report.

3 weeks later: review the last 5 earnings reports and explain their growth story.

2 months later: Here are 15 filings, earnings transcripts and research reports. Build an investment thesis.

Eventually: Monitor my stock portfolio and tell send me alerts when something materially changes my thesis.

Each stage requires different capabilities and hence different evals.

Your evals should change with usage patterns:

short-context -> long-context
single-turn QA -> multi-turn
passage citations -> doc and line citations
Simple QA -> complex synthesis
reactive chat -> proactive agent

If your evals are stuck in week 1 while users are in week 3, it will show in your product and churn metrics.

Here’s a practical way to build the roadmap:
1/ Map the dimensions along which usage will evolve (eg # turns, document size, tool use, autonomy, journey coverage).
2/ Prioritize the use cases and dimensions that matter most for your product.
3/ Talk to your users, mine production traces to look for shifts.
4/ Build P0 evals for the next stage of usage.
5/ Run the evals, find failure modes (remember the taxonomy from Part 3), and hill-climb (from part 6).

The goal is to stay ahead of usage patterns. PM 101.

Drop your questions and I will address in future posts.

Share this with your teammates who think your evals are perfect. [83 ❤️ 4 🔄]

Cat Wu (@_catwu)

  • Thanks to your feedback, we’ve unified memory across Chat and Cowork. Now, you can tell Claude to remember something once and it’ll have that context across surfaces!

Tell us what you’d like to see next :) [588 ❤️ 19 🔄]

Thariq (@trq212)

  • excited to share more on how we’re making Claude Code more hackable soon [1118 ❤️ 24 🔄]

Google Labs (@GoogleLabs)

  • 🚨 NEW LABS EXPERIMENT 🚨

Great minds t̵h̵i̵n̵k̵ play alike!

Play with Putty is a collaborative vibe coding tool that lets you build tools and websites together in real time. 🫟

Ready for your imagination to go multiplayer? Join the waitlist at and share your feedback with us!

US only, ages 18+ [1234 ❤️ 102 🔄]

Guillermo Rauch (@rauchg)

  • (Not affiliated with Denny’s) [31 ❤️]
  • We’re introducing Run SDK: secure 𝚎𝚟𝚊𝚕 for dynamic Code Mode execution.

When agents write code, you don’t always need a full sandbox. You can 𝚛𝚞𝚗 their code in a lightweight QuickJS secure context. Faster and more cost-efficient.

𝚗𝚙𝚖 𝚒 𝚛𝚞𝚗 [419 ❤️ 19 🔄]

  • The hardest problem in building agents is secure connectivity to services and data.

Excited for Vercel Connect to be GA. Run e.g: 𝚟𝚎𝚛𝚌𝚎𝚕 𝚌𝚘𝚗𝚗𝚎𝚌𝚝 𝚌𝚛𝚎𝚊𝚝𝚎 𝚗𝚘𝚝𝚒𝚘𝚗, then get an MCP client you can query on behalf of the authenticated user. [214 ❤️ 10 🔄]

Aaron Levie (@levie)

  • Good post on what the applied AI strategy looks like at scale. It’s clear that there’s a wide gap between the AI models and the underlying workflows of an enterprise, which leaves a ton of opportunity for applied AI companies.

“The world doesn’t just want raw models and agents; it wants problems resolved and outcomes achieved. The premium will sit with the companies that can diffuse this intelligence through every aspect of civilization, converting raw tokens into real world outcomes, transforming industries, and creating economies in the process.”

This requires understanding the context, driving the change management, having a harness that can route to various models, connecting to the critical business systems in that vertical, solving the UX challenges of connecting users to agents in the right way in a workflow, understanding the evals in the space, and so on.

That’s a ton of value that goes beyond just the model intelligence itself. And there’s a window of opportunity right now to build the defining companies that can bring intelligence to the critical domains in an enterprise. [122 ❤️ 6 🔄]

Garry Tan (@garrytan)

  • Ok apparently this guy came up with it

Don’t really care though

Clout chaser lmao he is the one who blocked me [60 ❤️]

  • lmao [4686 ❤️ 187 🔄]
  • You need to be cross-harness memory maxxing [510 ❤️ 28 🔄]

Nikunj Kothari (@nikunj)

  • find someone who loves you as much as the VC who loves posting photos of an “exclusive” dinner 🙈 [14 ❤️]
  • Originally inspired by the @TheStalwart’s Odd Lots episode.

Built using @ChatGPT codex and @Railway

Design polish courtesy of @emilkowalski skills and @itskasturiii generative loaders.

Direct link here:

Feedback always welcome! [3 ❤️]

  • Introducing the El Niño situation monitor (elneenyo dot com)..

real time updates, news and what’s happening straight from government sources
impact per region and costs
historical records and what’s the importance
glossary and FAQ on all the different readings

If this topic pops up in your group chat, you know where to go now! [8 ❤️]

Peter Steinberger (@steipete)

  • This week. [280 ❤️ 11 🔄]

Dan Shipper (@danshipper)

  • [37 ❤️ 1 🔄]
  • GET DOT [51 ❤️ 2 🔄]
  • I got one of these to test

@every vibe check forthcoming! [71 ❤️ 1 🔄]

Aditya Agarwal (@adityaag)

  • It’s totally unsurprising that the general population hates datacenter buildout.

AI today primarily helps knowledge workers and the highest paid segments of the country.

I suspect that the big change will happen when AI is able to find cures for diseases that ail everyone. We are starting to see light at the end of the tunnel here.

Also makes sense why some of the brightest minds now want to focus there.

Our industry also hasn’t done ourselves any favors by all our fear-mongering. Instead of painting a positive version of the future, we have instead talked about all the reasons why we should be afraid 🤦‍♂️. [18 ❤️ 1 🔄]

Sam Altman (@sama)

  • we made a chip and it is fast [32803 ❤️ 1347 🔄]

Claude (@claudeai)

  • Topics some consider sensitive, like health or religious beliefs, stay out of memory unless you turn them on in Settings.

Memory is on by default on Free, Pro, and Max plans. Review yours anytime in Settings > Memory.

Read more: [394 ❤️ 4 🔄]

  • Everything Claude remembers is saved as a list of topics in Settings, where you can read, edit, or delete each one.

Memory also updates on its own as you chat, saving new details as they come up. You can also say “remember this” to save something specific. [549 ❤️ 21 🔄]

  • Claude now has one memory across chat and Claude Cowork, and you decide what’s in it.

Hand Cowork a task and it starts from what Claude already knows from your chats: the project you talked through, your manager’s preferences, or the client from last quarter. [6846 ❤️ 399 🔄]


Follow Builders 自动生成 · 2026-08-26