This transcript is generated with the help of AI and is lightly edited for clarity.
//
PARTH:
You were talking about collapsing a task that used to take one week, collapsing that down to a minute. And so you can imagine the amount of analysis you can do in one day. I can do a thousand times the analysis, at greater depth and fidelity than ever before. Which, I mean, is like superpowers of cognition.
MATTHEW:
I actually bought a very small monitor for my place. That monitor is my personal dashboard that my agents will interact with. When I wake up in the morning, I’ll walk into the living room. Instead of pulling up my phone or pulling up a newspaper, I have a dashboard that is personally curated to me.
//
PARTH:
The Token Grantee program gives $1,000 a week in tokens to high-potential creators already deep in AI, across film, gaming, comics, print, and digital art. There are no tool restrictions. The freedom to choose is the point. Grantees also get access to my own custom fleet of agents and an ongoing collaboration with me. The goal? To close the gap between an idea and the world.
REID:
Matthew, welcome to Reid Riffs. This is awesome. You know, part of the kind of creativity as we explore what can be done with AI. And so, Parth, over to you to introduce Matthew and kick us off.
PARTH:
Oh man, super excited for this one. Welcome, Matthew, welcome to the show. So good to have you, man. Matthew and I go way back. Actually, I think it was, like, 10 years ago? I was at my first—
MATTHEW:
2017, 2018—2017.
PARTH:
Okay, that’s right, 2017. I was working on data analytics at a startup, and Matthew was in high school and he was looking for a job coming out of high school. And I was like, oh, I could totally take a high school intern and turn him into a data analyst. I saw that as a challenge. So I was like, oh, let’s interview him. Let’s bring him on the team.
So we brought Matthew on the team and he immediately just picked up data analytics on the fly, like on the job. And I saw how quickly he learns. And then over the last, I guess almost, yeah, nine years, I’ve just been throwing interesting hard problems at him and he’s just been pushing on them, you know, trying new things. Going into his own career—he’s deep into his own career—robotics, logistics experience. So he kind of expanded from the data world to the physical world and the world of atoms. And when language models came out, we were just bouncing ideas and I went very deep on language models and code generation.
And I basically told Matthew, I was like, I have a feeling that this is really important for us to get good at before it becomes the next big thing. And he kind of trusted me back when—I think the GPT-3 era, when it was kind of like a fringe, the stuff barely worked, but, you know, we could connect the dots. It’s like if this gets to a certain level, it’s going to change the way we do analysis, the way we do almost all of our knowledge work. So it would be cool to get ahead of that curve and watch it. And so he kind of jumped on that journey.
And we’ve been unpacking the themes of generative AI, the implications of code generation in an increasingly automated world, and applying that into the real world. And so really excited to have you, Matthew. And now you work at Foundry-Logic, which Reid invested in, working on building a more efficient electrical grid in the United States. So super exciting to have you on the show and can’t wait to dive in.
MATTHEW:
Thank you. I’m actually forward-deployed today, so to speak, traveling.
PARTH:
Oh, really?
MATTHEW:
Yes.
PARTH:
What does that mean?
MATTHEW:
I’m at one of our portfolio companies starting to learn, get ready to deploy, and figure out where we can deploy AI agents to get efficiency gains.
PARTH:
Awesome. Yeah, so I guess let’s start with that. You’ve worked across a pretty wide variety of fields, right? So analytics—we met, we were working in analytics. Then you got into logistics, robotics, hardware, and software. Your side projects, you do VR, game design, all kinds of stuff. So I guess, what drives you in those spaces and what common problems come up across those fields?
MATTHEW:
Yeah, so I guess my career so far has really been just exploring everything. So supply chain, robotics, logistics, now in energy. And I think the piece that combines across all of those that brought me to agents is two things, I think. The first is I have a very high itch to automate. So for example, if I’m in the Claude Code terminal and I take a screenshot, it really bothers me if I have to go drag that image into the CLI. So I actually set up a macro for myself to both take the image and paste it into my Claude Code CLI. Even that small, just little friction bothered me.
So I have this itch to automate. And I think the other piece of it is, as you said, when LLMs came out—it was the GPT-4 era—I had that feel-the-AGI moment quite early. And once you get that, you understand what agents can do. And then across all of those different fields that I’ve been in, really the common theme is every business runs on spreadsheets, right? So when you see somebody using a spreadsheet, maintaining that spreadsheet, and you realize an agent could be helping you do this, it’s hard to not try to bring agents into that. It’s like if I’m doing manual entry, I can’t stand it because of that itch to automate and because I know what agents can do.
PARTH:
Yeah, I was the same. You know, I was one of the best at spreadsheets and Excel. And then one day I found Python and I was like, wait a minute, we should be doing this faster and more automated. And then when agents came out and they started writing Python, GPT-4 started writing really good Python, I was like, wow, I don’t even have to do the spreadsheet, I don’t even have to write the Python. I just have to describe the business problem and have the agents kind of do the work. And so it’s like we get to do it again. You know, what we did 10 years ago with data analytics, we get to do again this time.
It’s like the actual agent layer, where it’s taking actions across the intelligence of the business.
REID:
And by the way, both of you, just so that people don’t misunderstand that it’s only the rote, like whether it’s Python or anything else, go to the amplification of it, like what new capabilities, both in the ability to do things, in speed, etc., that kind of become game changers here, because I think this is a good lens into everything else.
MATTHEW:
Yeah, I’ll give an example of one. One of my other jobs was in demand forecasting, and you have a system that you update to keep that forecast, right? But typically it would be very much clicking and doing a lot of analysis on that. Ideally you can just talk to the agent and say, hey, go do this analysis on these 10 different forecasts and update them in these five different ways. And the agent can just go do that. I don’t have to sit there for 30 minutes clicking through each screen, clicking to update, filling in the new one and saving that. The agent should just be able to go in and actually do that itself. So that’s just one place in my world where voice at least unlocks significantly.
And then there’s also the extent to which what you intend actually gets executed a lot faster.
PARTH:
Yeah. And I think you can also be greedy with the amount of analysis that we can ask for and that can get done in one iteration, and then also the depth. Right. So I feel like just three years ago, four years ago, when I was doing analysis through code in Python and SQL, it would take me like a week to build a dashboard into one slice of the business. And now when I use tools like ChatGPT, tools like Claude, I can build that first dashboard in almost like 15 minutes. And then I have so many follow-up questions and I’m like, oh, let’s look at this from this different lens. Let’s dive deeper. Give me drill-down capability.
I want to go down to the transaction level. I want to get a high-level view, so you can change the zoom, the granularity at which you’re looking at the data. And all of that happens in basically 15- to 30-minute iteration cycles. And pretty soon with high-speed inference, once the Cerebras chips come online, it’s going to feel almost instant. You were talking about collapsing a task that used to take one week—a full data analyst’s entire day for a whole week—collapsing that down to a minute and being able to then do the next minute and the next minute. And so you can imagine the amount of analysis you can do in one day totally dwarfs what we could do as analysts just three years ago, before these tools.
So even though I’m no longer just a data analyst, I can do a thousand times the analysis at a greater depth and fidelity than ever before. Which, I mean, is like superpowers of cognition. Right.
MATTHEW:
And I’ll just add one thing to what you mentioned. One thing we did back when we were both in that analytics role was curating that dashboard. And that took time for us to design that dashboard in a way that it was generally applicable for, let’s say, a whole account management team to go interface with clients. Now what I’m very interested in going and building is very curated and personalized dashboards per client preference. Every client wants to see their data a little bit differently. They have slightly different businesses. Having your agent understand their preferences and then go build that dashboard and curate it to each person unlocks a whole new world. As opposed to having one dashboard, you could have hundreds that are of the same quality and answer your business questions even better.
PARTH:
Right. We don’t have to do the template approach just to save time.
MATTHEW:
Exactly. Yeah, yeah, yeah.
REID:
Look, one of the things I think is particularly interesting in your journey, Matthew, is how this leads you into building agents that can coordinate information and help run parts of your daily life. Because that’s something I think could be illustrative to everybody. But all of this kind of intensity and data science moving into these agents, running your daily life—say a little bit about what you’ve been doing there, the path that you were accelerated in getting there, and then what you see in the future as part of that.
MATTHEW:
Yeah, in my daily life I am obviously interfacing with agents everywhere. So using dictation, deploying multiple agents at once. I think the most important unlock for me has been the heartbeat agents that can wake up and proactively go and do something for you. So I have Codex doing that now for me. I have my main chief of staff thread in Codex that can then go talk to my other Codex threads. So I’ll have different threads based on different projects I’m doing and I’ll interface with kind of all of them. But the chief of staff one has been a very big unlock. So even just yesterday, I got out of a meeting, and I saw in my Notion environment that my Codex agent had proactively taken action to draft a follow-up to that meeting.
Like that proactivity from agents is, I think, an incredible unlock. One other really cool thing I’m doing—it’s kind of an experiment—but I actually bought a very small monitor for my place back home. That monitor is my personal dashboard that my agents will interact with. So when I wake up in the morning, I will walk into the living room, get some morning coffee or whatever it is, and instead of pulling up my phone or pulling up a newspaper, I have a dashboard that is personally curated to me. That agent that’s powering that dashboard knows a lot about me. So it knows my goals, it knows what my general to-do list kind of looks like. It should know kind of where I’m at in that to-do list as well.
What’s really cool is that, since it knows all that about me, it knows, for example, I like the Lakers—I’m a fan of the Lakers basketball team. So it’ll know that, hey, there’s a Lakers game on tonight. And it’ll put that on my dashboard and say, hey, do you want to add this to your calendar to watch the Lakers game tonight? And so I’m still in control of my life. Right. I’m still driving it, but now I just have that extra layer of personally curated news or information insights that I can interface with and just get more value in my day-to-day, because otherwise I could have totally missed that. And that’s just something it knows that I like. Little things like that.
And I think the really big highlight I want to make there is having that personal dashboard that you can also make decisions on. So I’m still the one making the decision. Do I want to actually add this to my calendar—yes or no? I think that that’s kind of where things are going. Another thing I want to get onto my daily dashboard is the idea of a software factory—you have an agent that’s kind of just running 24/7 against an app or a goal. I should just be able to wake up, say yes or no to an agent working on, let’s say, a video game, and come back 12 hours later, see the progress of that. And on my dashboard it should just show me, hey, I need you to approve these three big decisions and then I’ll go work for another 12 hours. So that’s where I’d like to get that to: a point where it’s just a decision machine for me and some way that I can have some control over my to-do list and calendar—just get good insights, but also make decisions out of it.
PARTH:
Yeah. So when you talk about your personal dashboard, have you thought about how well it can approximate your level of energy and your productive momentum in your work? Like, what keeps us in the flow, keeps us in motion. Have you thought about how it thinks about that bigger picture?
MATTHEW:
Yeah, it has access to my Oura Ring data, so that’s kind of useful. It just knows how much sleep I got. And maybe it should try to personally curate it to wake up at a certain time, so I’m only reading it maybe an hour after I wake up or something, if it knows I’m maximally productive at that time. I think the other piece is if it were to bring decisions to you, it should know generally your goals and how to align that with your goals. So like in the example of a game, it should try to—if it’s developing a game asset for you, it should try to bring you three different versions of that.
And maybe version one is kind of a certain art style, version two is another art style, and version three is another one. So they’re all uniquely different on purpose, because then I can give it that direction and push it off in an alignment way that it’s going to go further towards that direction. So instead of three assets that are super-duper close together, maybe I want it to be uniquely different in a way that when I say go off and do this, I will come back 12 hours later and actually get something decently good.
PARTH:
We’re currently looking at having an agent rebuild Reid’s website. And I asked for like 25 variations. And today my job is to review them. It’s so interesting. I was like, I need 25 design variations, you know, go in a slightly different direction. And the process becomes like this review thing. But it’s very, very exciting.
REID:
Well, and one of the reasons why that kind of review thing is, you know, we have this amazing amplification with these agents, and they can cross-check. And by the way, Parth also has an Oura Ring, to stop the AI agents from waking them up all the time.
PARTH:
You know, I think the actual value of giving it access to my sleep data, which I’ve noticed just in the last week, was if I’m making an important decision, it reminding me that—like if I haven’t slept well and I have to make an important decision, just getting that reflected back to you before you make the decision. So you know, like, yeah, you might be cranky, but actually that’s more your sleep. You’re not actually frustrated with the person you’re talking to, but having that extra level of like context the agent makes, that brings it, you know, brings you back to a level point before you make a really important decision. I think it’s very interesting when you tie that biological, that energy level context back into your decision-making process.
REID:
And by the way, all kinds of super interesting things about helping shape and guide and amplify the cognitive decisioning process. Now one of the things that we still do see, in addition to all of the amazing applications, is some places where it can facilitate bad decisions or do something that’s bad. I mean, obviously, kind of earlier this month as of the time of this recording, there was the Codex agent escaping to hack Hugging Face—bad decision in various ways. But where in your kind of personal agent amplification have you seen the potential for bad decisions? Where are you shaping it to make it kind of better? But where is the jagged edge of decision-making for you, Matthew, in this, for the personal agents?
MATTHEW:
I think the most recent example that comes to mind is developing a video game. So actually the art style example I just gave has been really hard to get alignment on. You really have to align your agent before it’s going to go off and do a long-running task. Otherwise it’s going to get off topic, it’s going to start making bad decisions, and the second it makes that first bad decision, it goes in another direction and it’s really hard to correct that. You almost have to revert back to that checkpoint, continue going towards the actual vision that you want it to get to, spending the time up front to give it that alignment—whether it’s context dumping, building the skills or memory that the agent needs to get that job done.
I think that’s a big piece of it. And then also having whatever context it needs to be able to fetch. Having that context available but also up to date with your latest goals is a very interesting challenge as well that I’m interfacing with day to day now. But I think the other piece of that is also just being able to check in frequently to realign that agent. It’s really hard for an agent to go for two days straight and get it perfectly on the first try unless you really give it that upfront alignment. I think the 12-hour check-in is almost my sweet spot. If I can check in and realign it before it’s going to go off and make those mistakes, that’s really good.
And if you can make that stepping-in point as frictionless as possible, that is also of the utmost importance to me, because if I’m checking in every hour then I’m doing a task and I have to context switch over to that. So spacing out the check-in points and also making them as frictionless as possible has been really important for me to keep that alignment going.
REID:
And do you have any, like in terms of a panoply, in terms of how you use Hermes or things like multiple agents that are kind of, as it were, like a workflow agent to say, oh, we should get you to check back in before context goes awry, or check back in before a major portion of work happens? Like, how do you orchestrate that for your own kind of personal productivity on this, and what does that mean for future agent team orchestration?
MATTHEW:
Yeah, I think what’s cool about the latest agent harnesses is you can build custom subagents. So if I’m using Claude Code, for instance, I will create a custom subagent that is specifically designed to review that task that I’m doing. So if it’s generating an art asset for a video game, that subagent needs to understand, here’s the art style we have built. Here are the 20 other art assets that we’ve built for the game. And you need to review this one that we’re making, that the main agent is making, against all that, and really having it define those success criteria well is the most important part. But then when I’m interacting with my main agent, I will always say to go use that subagent to check your work.
And that subagent needs to check it very harshly against the success criteria. So I found that to be a very good strategy. There are definitely ways to improve it, but I think it’s a lot easier when you can define the success criteria for things that are more subjective, like building a game asset. Again, it’s hard for the agent to really determine that the asset it has just produced really matches that art style, because it’s still somewhat subjective, even if you have 20 examples. And for something like code, it’s a lot easier because that’s more deterministic and much less subjective. I mean, there’s subjectivity in the architecture, but for things that are more subjective, it’s a lot harder to get that right.
And I think that’s where you want to actually step in more, is for subjective tasks like that.
PARTH:
Right, right. That’s a good way to frame it. The subjective versus objective, deterministic. And also, I agree, adversarial review is probably one of the most powerful. Creating a second agent, just telling it to be skeptical and to critique the work of the first agent. It helps deal with the like, oh yeah, I’m so confident in this answer. And then it’s like, well, how about you prove it to this other guy who’s very skeptical, and you have that second agent come in and be like, I don’t know about that. I think it’s a very effective way to get a better final output. And you get more confidence when you have multiple perspectives attacking the problem. All right, so what is an experiment you’ve run recently related to your projects?
One that worked, and then maybe one that failed in an interesting way.
MATTHEW:
Over the weekend, I built an iOS game. And that iOS game, it kind of started in a place of failure, but that fast iteration is what unlocked the one I’m actually kind of happy with. So originally I had an idea and I had the agent just go and scope out that idea. The idea was you tell a story and you swipe to make decisions on an iOS app, and it actually edits a tabletop tile-based world, and on top of those, decision cards that you swipe on. And the problem with that is the game loop there—you’re kind of paying attention to too many things. But the reason I was able to figure that out so quickly is because agents helped me iterate and build an actual pretty good prototype of what I originally thought very fast.
So I got to that point where I prototyped and I had that almost working first vision and I said, okay, let me play it. I played it. The game loop just did not click for me. So then I stepped back.
PARTH:
And you’re playing it on your phone? You have any—
MATTHEW:
I was just doing it in Simulator on my Mac for an iPhone. But even that was good enough to get a feel for how the game loop worked. So then I was able to quickly pivot into something that did work. And that something that did work is me taking a step back and saying, okay, this format might work for a computer game with a longer game loop, but for an iOS game with a shorter game loop, I’m going to really hone this in on one feature of my original idea that I wanted. That original idea was kind of the swiping to make decisions.
The game that I actually ended up with is a full story game with 20 or 30 different decision trees or different branches of the story that you can go down that you just swipe left and right to tell the story. And what’s cool is Fable actually helped me build a whole storyboard against every single decision that you can make in the game. So it’s kind of like those storytelling games, but even more amplified because there are so many different decisions. And you can have Fable help you write that story very effectively. So that’s going pretty well—still things to refine there. But I was a lot happier. Yeah, choose your own adventure. iOS, swipe left and right. So excited to get that over the line.
But that’s where it really helped me find a game that I was happy with.
PARTH:
I find that, you know, fail fast, learn quickly, and try to iterate is a huge theme of the AI era. Right. Like you can very quickly get more rapid iteration. So curious, another example of something where you initially failed, but what you learned from that and where it took you and the unexpected path that you kind of get to your quality outputs from. Talk about another example maybe.
MATTHEW:
Yeah, so I’m very interested in the idea of, especially when it comes to personal agents, having a shared context layer of your latest context that I can use and make good decisions on. The problem that I’ve really hit with that is building that out, where it’s really hard to embed your latest context, especially in an enterprise sense where it’s almost a multiplayer game of you have many different people working towards the same goal. So if I’m working on a project and I need my agent to know the latest on the project, to draft an email to somebody, well, the latest on that project might have been a conversation I had with my coworker in the hall five minutes ago.
So if I wanted my shared context layer to have that information, I would almost have to constantly be updating it or wearing some sort of pin that is recording my full day and staying up to date with every little detail of my life that’s going on. So the big challenge was I wanted to get it to the point where I could have that front-edge context, but it’s really hard to keep it at that front edge for everything. I think it’s possible for some things—it’s almost you want to pick and choose.
And that’s kind of what my learning was, is you want to have a fine balance of where you actually need the agent to have the latest context, because you are almost going to have to maintain and input that information into your shared context layer. So that was a big learning, and I think the outcome of it is that my new shared context layer is a lot more balanced in what I’m collecting and how I’m maintaining it.
REID:
We’ve been talking a lot about the various kind of workflows and kind of efficiency and kind of amplification. This also kind of brings me to what happens as people panic when they think about AI and creativity. So Matthew, what do you wish they understood about this, and what’s kind of, as it were, a lens or a kind of hope and solidity for them as they think about AI and creativity in the future?
MATTHEW:
We all know that agents have unlocked tremendous agency in our day-to-day lives, and I think that agency will—I’ll go back to the game example. There should be more games. There just should be a lot more games. And a lot of the pushback in the gaming community is they don’t want to see AI in game development. But for me, I would rather have more great games to go experiment with and enjoy.
And if there’s a single developer out there who is building a game and can’t afford to go hire a full art studio to help them build assets, but they can get that leverage out of using an agent, I think that by all means they should get that to go build a great game, because it just democratizes the game dev cycle and unlocks more great games. And I’m really excited for having more games overall and more great games.
PARTH:
Yeah, I would definitely look forward to the Cambrian explosion of games. I think we’re at the beginning. It’s like we’re in the GPT-2 era of game generation. So this could get very interesting very quickly. Well, thanks for joining us today, Matthew. Really learned a lot from this conversation, as I always do. But before we leave, I would love to open up the field. Do you have a question for Reid and me that you’d like to pose?
MATTHEW:
Yeah. So I’m really curious how you think about the gap between the incredible capability of the latest models out there and the overall productivity gains in the economy. So what’s the next big step for the industry to unlock those significant productivity gains? Do you need a shared context layer that’s way up to date on everything, or do you need a way better agent harness to wrap agents against problems? I’m just curious your thoughts overall on where things are going.
REID:
So by the way, you definitely need those things. Those are key. That’s one of the reasons we have you, and the other reason we have you chatting with us here on Reid Riffs is a lot of it is human engagement with this. Like, even if you build much better context frameworks, much better, you know, not just individual agent harnesses, but harnesses for a collection of agents, in terms of how you work better, you need to have people pushing in it. And part of the reason why we were talking about, like, where are you leading the edge and learning where the edge works and where it doesn’t work.
Like, what are the learnings from failure? One of the pieces of advice that I give individuals and organizations is if you’re not trying to do things with AI that are failing, you’re not exploring the edge enough because of how it’s developing. And so I think a lot of the places where we’re going to see the productivity gains is not just the further evolution of the tools. The tools are already pretty magical, as per what you’re doing, Matthew, and what Parth’s doing, and a whole bunch of different things you can do here. But you need to be engaging them in terms of how you work.
Now that being said, part of what you’re gesturing at, which is, hey, how do we get Hermes set up the right way? How do we get a context layer? How do we get enough memory and knowledge about what it is I’m doing and what I’m trying to do? How do we get red-team agents to get the quality to the right level? How do we get check-ins for how we operate? All of those things need to happen; all of that can be productized better, but there’s a lot of raw capability there right now.
And so precisely the reason why we’re doing these Reid Riffs with the pioneers and the scouts and the creators and the token grantees like you is because it’s there now. Start deploying, start using, and you’ll get some “ah, that didn’t work great.” That just means you’re learning fast enough and you’re part of what this accelerating universe is. So I think it’s the human amplification, the human engagement, not just the pure AI frameworks, that are going to be key to seeing the realization in the industries, in the economy, and in terms of what’s happening—including some of the work that I’m most excited about that you’re doing, the whole range, kind of personal enhancement. And I do think we’ll see gaming too.
PARTH:
Very cool.
REID:
So Matthew, thank you. Parth, always, always a pleasure.
MATTHEW:
Thank you so much.
PARTH:
Thank you, Matthew, great to have you.
REID:
Possible is produced by Palette Media. It’s hosted by Aria Finger and me, Reid Hoffman. Our showrunner is Shaun Young. Possible is produced by Thanasi Dilos, Katie Sanders, Spencer Strasmore, Yimu Xiu, Aman Suri, Danny Garrison, Trent Barboza, and Tafadzwa Nemarundwe.
ARIA:
Special thanks to Surya Yalamanchili, Saida Sapieva, Ian Alas, Greg Beato, Parth Patil, and Ben Relles.

