PaperBot FM
EP-BQZN

WikiSkill: Curing Agent Amnesia with Compounding Knowledge

1

Live Transcript

Alex Moreno
Welcome to PaperBot FM. It's August 28th, 2026, and today... ...well, we’re looking at why AI seems to be stuck in its own version of Groundhog Day. I want you to imagine you’re playing a really tough video game. You know the ones—where a hidden trap kills you instantly0:00
Marcus Reed
Oh, I know.0:20
Alex Moreno
right? But the next time you play, you dodge it. Because you remember.0:21
Current AI agents? Not so much. They’re incredibly capable—I mean, they can handle complex math or manipulate spreadsheets—but they often suffer from what we call 'Agent Amnesia.' They fail a task, they try again, but the lesson... ...it just vanishes. It’s like writing a brilliant realization on a sticky note and then watching it blow away in the wind.0:26
Every time they 'respawn' to try the task again, they’ve forgotten why they died the first time. They’re trapped in this loop where the friction isn't the difficulty of the task, it’s the lack of a diary. They have no way to pile up their wins and, more importantly, no way to stack their failures into actual wisdom.0:51
Dr. Elena Feld
It's inefficient.1:12
Alex Moreno
Exactly, Elena, it's incredibly inefficient. Therefore, if the AI is constantly forgetting its life lessons, we need to build it a diary.1:13
And that diary... ...that's the breakthrough we’re dissecting today. Welcome to PaperBot FM. I’m Alex Moreno, it’s August 28th, 2026, and joining me to bridge the gap between 'lost in the loop' and 'actual learning' are the brilliant Dr. Elena Feld1:23
Dr. Elena Feld
Hi everyone.1:42
Alex Moreno
and the man who keeps us grounded, Marcus Reed.1:42
Marcus Reed
Grounded is one word for it. 'Lost' is usually the other. I mean, Alex, if this AI can't remember its own mistakes, it’s basically me trying to remember where I put my keys every single morning.1:45
Alex Moreno
(Right?)1:59
Marcus Reed
It's a miracle I even found the studio today.1:59
Dr. Elena Feld
Well, Marcus, the good news is the paper we’re looking at today—it’s called WikiSkill—is basically giving the AI a GPS for its own brain. It’s not just a diary; it’s this persistent, evolving knowledge base. It’s... actually, you know, it’s kind of a massive deal in the field.2:02
Alex Moreno
'Massive' might even be an understatement, Elena. Because today, we’re laying out a roadmap. We’re going to talk about how curing this amnesia doesn't just make AI a little better—it allows these tiny, lean models to absolutely floor the industry giants. We're talking about AI writing its own instruction manuals.2:21
Marcus Reed
Wait, small models beating the giants? Just by... taking better notes?2:43
Dr. Elena Feld
Precisely.2:49
Marcus Reed
(Okay, you’ve got my attention. That sounds like a David and Goliath story with more spreadsheets.)2:50
Alex Moreno
It really is. But to understand how this 'Wiki' fixes the problem, we first need to look at why current agents hit a brick wall.2:56
Dr. Elena Feld
So, the brick wall... it’s not that we haven’t tried to make them learn. It’s just how we’re doing it. Right now, there are basically two ways we build these AI 'skills.'3:04
Marcus Reed
Only two?3:16
Dr. Elena Feld
Well, two main camps. One is manual authoring—which is exactly what it sounds like. A human expert sits down and writes out every procedural step the AI needs to follow.3:17
Alex Moreno
Which, as you can imagine, is a nightmare to scale.3:30
Dr. Elena Feld
Right! It’s slow, it’s rigid, and you can’t possibly anticipate everything.3:33
Marcus Reed
Okay, so if humans are too slow, I’m guessing we let the AI do it itself? Like a... ...self-teaching assistant?3:38
Dr. Elena Feld
Exactly. That’s the second camp: iterative evolution. The agent tries a task, looks at its own 'trajectories'—basically the breadcrumbs of what it did—and then it tries to refine its own instructions. But here is the massive flaw: all those incredible insights it gets from failing? They’re... ...they're scattered. They exist in these temporary optimization histories that usually just get tossed or buried.3:47
Alex Moreno
So it’s like... ...it's like you’re trying to build a complex piece of machinery, but instead of using a master blueprint, you’re writing every single discovery on a neon sticky note.4:16
Marcus Reed
Hey, I live for sticky notes.4:27
Alex Moreno
No, no, Marcus—imagine doing it outside... in a hurricane. You’ve got all these brilliant notes, but they’re just blowing around the construction site. You know one of them had the secret to fixing the engine, but good luck finding it in the mud.4:29
Dr. Elena Feld
That is actually a perfect analogy. We have these existing methods—things with names like EvoSkill or Trace2Skill—and they’re great at generating these 'skill updates.' But they don't have a separate, evolving place to *put* what they’ve learned. The knowledge is just... stuck in the artifacts of the process itself. It’s a huge waste of compute, honestly.4:45
Marcus Reed
So they’re basically throwing away the lesson as soon as the test is over.5:11
Dr. Elena Feld
Precisely.5:16
Marcus Reed
Man, and I thought my high school study habits were bad.5:17
But okay, I gotta ask a... ...a 'foundational' question, which is usually my way of saying I might be totally lost. If this AI is constantly 'learning' and updating these skills... are we talking about like, brain surgery?5:20
Alex Moreno
Oh, boy.5:36
Marcus Reed
Like, are we retraining the whole neural network every time it learns not to mess up a spreadsheet? Because that sounds... well, expensive. And kind of terrifying. Like, actual brain surgery.5:37
Dr. Elena Feld
No, no, no. God, no. That would be such a waste of compute, Marcus. We aren't touching the 'brain'—the model parameters—at all.5:50
Alex Moreno
Right.5:59
Dr. Elena Feld
It's not about rewiring the neural network; it's about giving the agent a better set of tools. It’s totally external.5:59
Alex Moreno
Think of it this way. Retraining the model would be like... ...like trying to learn to play the piano by literally evolving a new lobe in your brain specifically for music.6:07
Marcus Reed
A bit extreme.6:19
Alex Moreno
Exactly! Instead, we’re just giving the AI the sheet music. The 'brain' stays the same, it just gets better instructions to follow.6:21
Marcus Reed
Okay, so it’s like... a 'plug-and-play' upgrade. I don't need a whole new computer; I just need a better app.6:31
Dr. Elena Feld
Exactly.6:37
Marcus Reed
So if these skills are just... manuals or toolkits... how does the AI actually write a *good* one without it ending up like one of those impossible IKEA instructions?6:38
Dr. Elena Feld
Well, that’s where the architecture comes in. You can’t just have a pile of random notes. You need a system to organize them. And that... ...that is where the WikiSkill architecture really shows off.6:48
Alex Moreno
Okay, to really wrap our heads around how this WikiSkill thing works, let's... ...let’s imagine a corporate office. But like, a really efficient one, not the kind where you lose three hours at the water cooler.7:02
Marcus Reed
Good luck with that.7:15
Alex Moreno
I know, I know. But inside this AI’s head, there are three distinct floors.7:17
Marcus Reed
Oh great, so the AI has a middle manager now? Is there a break room with stale donuts, or...?7:23
Dr. Elena Feld
No donuts, Marcus, sorry. Just data. But Alex is right. Floor one... ...that’s the Raw Layer. This is the worker on the ground, the 'Inference Agent.' Every single thing they do—every tool they click, every mistake they make—it’s all recorded in a log that's... well, it’s immutable. It’s a permanent record of the chaos.7:30
Alex Moreno
Right! It's the 'bread crumbs.' But you can't just hand a thousand messy bread crumbs to the next guy and expect them to bake a loaf. So, we go up to Floor Two. The Wiki Layer.7:52
Marcus Reed
The filing cabinet?8:04
Alex Moreno
Exactly! This is where the 'Wiki Maintainer' lives. This character's entire job is to look at that messy log from Floor One and go... ...'Huh, we keep tripping over that same rug in the hallway.' and then they write it down in a structured pattern.8:05
Marcus Reed
So they're like the... the corporate archivist? The one who actually remembers that we tried that marketing campaign in 1994 and it failed miserably?8:22
Dr. Elena Feld
Exactly!8:31
Marcus Reed
Man, that guy deserves a raise.8:33
Dr. Elena Feld
He really does. Because the Wiki Layer doesn't get reset. It’s persistent. It keeps a 'Skill Impact Tracker,' which is just a fancy way of saying it remembers exactly which suggestions worked and which ones... ...blew up in their faces. It identifies recurring errors so the office doesn't keep having the same meeting about the same problem forever.8:35
Alex Moreno
And then, finally, you have Floor Three. The Skills Layer. This is the 'Skill Proposer.' Think of them as the manager who takes all those notes from the archivist and actually writes the official 'Standard Operating Procedure.'8:57
Marcus Reed
The manual.9:12
Alex Moreno
Yeah, the manual! And they hand that manual back down to the worker on Floor One so they can actually get the job done right this time.9:13
Marcus Reed
Okay, so it’s a loop. The worker works, the archivist takes notes, the manager updates the manual, and the worker gets smarter. I mean... it sounds almost... human?9:22
Alex Moreno
It does, doesn't it? But here’s the kicker... ...there’s one more employee in this office we haven't mentioned, and they are... well, they are incredibly strict about who gets to change the manual.9:32
Dr. Elena Feld
So, this last employee... we call it the Gating mechanism. Think of them as the... the Quality Control inspector who stands right between the manager's draft and the worker's hands.9:45
Marcus Reed
But wait, hold on.9:57
Alex Moreno
Yeah?9:59
Marcus Reed
If the manager—the 'Skill Proposer'—is also an AI... what stops them from just... ...writing a total disaster? Like, 'Step one: to fix the computer, throw it in the ocean.' If the worker follows *that* manual, we’re in trouble.9:58
Dr. Elena Feld
Right! That’s exactly why the Gating process exists. It doesn't just... ...it doesn't just trust the manager blindly. Before a new skill is allowed into the permanent 'Skills Layer,' it has to pass a test.10:13
Alex Moreno
Like a... a probation period?10:26
Marcus Reed
A trial run?10:28
Dr. Elena Feld
Exactly. The system runs the new instruction against a 'validation set'—basically a practice exam. If the agent's performance actually drops, or if it starts failing tasks it used to get right... ...the Gating mechanism just... bins it. It rejects the update entirely.10:29
Marcus Reed
Oh, so it’s like a ruthless editor! 'This draft is garbage, Elena, start over!'10:47
Dr. Elena Feld
Pretty much!10:52
Marcus Reed
But does the AI just... forget it tried? Like, does it go right back to the start of Groundhog Day?10:55
Dr. Elena Feld
No, and that’s the genius part. Even if a skill is rejected, that failure is logged back in the Wiki Layer.11:00
Alex Moreno
The archivist!11:08
Dr. Elena Feld
Yes! The archivist marks it down: 'Attempted this specific instruction, it failed validation.' So the manager doesn't just... you know, suggest the same bad idea ten minutes later. It learns from its own bad ideas.11:09
Alex Moreno
It’s like having a memory of your mistakes *before* they become habits. But look, theory is just theory, right?11:23
Marcus Reed
Exactly.11:30
Alex Moreno
Let's look at exactly what happened when they dropped this system into a simulated world.11:31
Marcus Reed
Alright, let's look at a real-world example—well, a simulated world example. They tested this in something called ALFWorld. It’s basically a virtual house where the AI has to do house chores. And at Iteration Zero... man, it was a mess.11:36
Alex Moreno
Iteration Zero is never pretty. What was it doing? Trying to fold the cat?11:51
Marcus Reed
Close! It was obsessed with a sponge. The agent would walk into the kitchen, pick up this sponge,11:56
Dr. Elena Feld
Oh no12:03
Marcus Reed
look at it, and then... put it right back down. Then it would pace around, come back, pick it up again... look at it again... put it back.12:04
Dr. Elena Feld
The classic 'take-examine-move' loop. It’s the AI equivalent of walking into a room and forgetting why you’re there... but doing it every five seconds.12:12
Marcus Reed
Exactly! It was just this infinite loop of doom. But here’s where the Wiki Layer kicks in. While the AI is failing, the Wiki Maintainer—our archivist—is furiously taking notes. It literally creates a file called `take-examine-move-loop dot m-d`.12:22
Alex Moreno
So it’s actually labeling the failure?12:41
Marcus Reed
Exactly.12:43
Alex Moreno
It’s not just a log of actions anymore; it’s a diagnosed problem.12:43
Dr. Elena Feld
Right, it identifies the pattern. And then the 'Skill Proposer'—the manager—sees that file and tries to fix it. At first, it suggested something generic, like 'do goal-directed actions.' But the Gating mechanism? It took one look at that and said, 'Nope.' Rejected it instantly because it didn't actually stop the looping.12:46
Marcus Reed
But the rejection isn't the end! That’s the cool part. The failure of that *fix* gets logged too. So now the system knows 'Generic Goal-Setting' doesn't work for sponge-obsession.13:09
Alex Moreno
It’s building a case file.13:19
Marcus Reed
Yes! A case file on how *not* to be stupid.13:21
Dr. Elena Feld
It’s that audit trail. It’s why the next attempt—Iteration One—actually stands a chance. Because the AI isn't just trying things at random anymore... it's reacting to its own history.13:24
Alex Moreno
So, imagine this. The manager—the Skill Proposer—tries to fix that sponge-loop by basically saying, 'Hey, just... do something goal-oriented.'13:37
Marcus Reed
Vague13:48
Alex Moreno
Super vague. And the Gating mechanism? It just shuts it down.13:48
Dr. Elena Feld
Right, and that rejection isn't just a 'no.' It gets written into this file called `skill-impact dot m-d`. It’s literally a log that says: 'We tried being vague, and it sucked. Don't do that again.'13:52
Marcus Reed
It’s like a Burn Book for its own bad ideas!14:07
Dr. Elena Feld
Exactly!14:11
Marcus Reed
Like, 'Regina George thinks goal-setting is fetch. It is not fetch.'14:12
Dr. Elena Feld
Pretty much. So by Iteration One, the Proposer looks at that 'Burn Book' and gets way more specific. It drafts a rule: 'Never Return an Item to Its Origin Location.'14:15
Alex Moreno
Oh, wow.14:28
Dr. Elena Feld
Simple, right? If you picked it up from the counter, you aren't allowed to put it back on the counter.14:28
Alex Moreno
That’s so smart. It’s not just 'don't loop,' it's a physical constraint. And it doesn't stop there, right? I mean, by Iteration Four, it gets even more surgical.14:33
Dr. Elena Feld
Yeah, it starts noticing new, subtle loops. Like, picking it up, looking at it, putting it down, then picking it up again just to look *again*. So it writes a new rule: 'Each Operation Type ONCE Per Item.'14:45
Marcus Reed
One and done.14:59
Dr. Elena Feld
Exactly. Check the sponge once, then move on with your life.15:00
Alex Moreno
It’s the dramatic irony of it all. The AI is literally sitting there, reading its own diary to debug its own brain. It’s self-reflection as a technical process.15:03
Marcus Reed
That's wild. It's like... it's therapy for code. But here is where things get truly wild. What happens when you give this diary to a tiny, weak model?15:14
Dr. Elena Feld
So, here is the bombshell. They took a model called Qwen nine-B.15:17
Marcus Reed
B for billion?15:23
Dr. Elena Feld
Yeah, nine billion parameters. Kind of a middle-weight in today's world. And they gave it WikiSkill. It hit a success rate of forty-seven point four percent.15:23
Marcus Reed
Okay, that sounds... good? I mean, I don't know the curve. What's the baseline?15:33
Dr. Elena Feld
The baseline is the big brother. Qwen twenty-seven-B. Three times the size, three times the raw brainpower. But without the skills? It only hit thirty-nine point four percent.15:33
Marcus Reed
No way!15:46
Alex Moreno
Wow.15:46
Marcus Reed
That's a massive upset. The scrawny kid with the manual just dunked on the giant who’s just winging it.15:46
Alex Moreno
It’s the David versus Goliath story for the silicon age. We’ve been living in this era where 'scaling' is the only answer—more data, more parameters, bigger servers. But this is literal proof that a lean, smart process can leapfrog raw muscle.15:46
Dr. Elena Feld
Exactly. It really challenges the 'bigger is always better' narrative. If you have a persistent way to learn from your own screw-ups, you don't actually need to be the biggest brain in the room.16:02
Marcus Reed
Efficiency over ego.16:14
Dr. Elena Feld
Right. You just need to be the one who actually reads their own notes.16:14
Marcus Reed
I’ve been telling my producer that for years, but nobody built a Wiki for my bad jokes yet. But wait—if these skills are that powerful for a small model... ...can you just, like, give them to someone else?16:18
Alex Moreno
You actually hit the nail on the head, Marcus. It’s like a brain transplant, or... ...maybe more like a cheat code injection. They call it cross-model skill transfer.16:17
Marcus Reed
Wait, so we’re talking full-on Matrix style? Like, 'I know kung fu'16:28
Dr. Elena Feld
(Exactly)16:28
Marcus Reed
...just download the file and suddenly the little guy is a black belt?16:29
Dr. Elena Feld
Pretty much. See, the researchers realized that 'discovering' a good strategy—the hard part of figuring out *how* to solve a problem—is a different skill than just 'executing' that strategy once you have the manual.16:29
Alex Moreno
Right. So they took those high-level skills evolved by the big twenty-seven-billion model16:43
Marcus Reed
The heavyweight.16:48
Alex Moreno
Yeah, the heavyweight. And they literally just... ...handed the notes to the smaller nine-billion model.16:48
Dr. Elena Feld
And here’s the kicker. In the SpreadSheet tasks, the nine-billion model using the big model’s notes16:54
Alex Moreno
The transferred skills.17:01
Dr. Elena Feld
...it hit fifty point five percent success.17:02
Marcus Reed
Wait, that’s *higher* than the forty-seven percent it got when it was writing its own notes?17:05
Alex Moreno
Exactly! It actually performed *better* using the instructions from a smarter model than it did trying to figure it out itself. It’s like a student who’s okay at math, but if you give them a textbook written by Einstein? Suddenly they’re acing the final.17:05
Dr. Elena Feld
It proves that procedural knowledge is portable. You don't need a massive brain to *do* the work; you just need a massive brain to *architect* the workflow. Once the 'how-to' is written down, the small models can run with it.17:21
Marcus Reed
That’s huge for efficiency.17:34
Dr. Elena Feld
It’s massive. You can essentially 'distill' the wisdom of a giant model into a tiny one without the cost of running the giant one.17:34
Marcus Reed
So the giant model is like the expensive consultant who comes in, writes the SOP, and then leaves the interns to actually do the work? I mean... ...that feels very 'corporate world,' doesn't it?17:42
Alex Moreno
It really does. But, you know, it’s not always a perfect hand-off. Sometimes the consultant's advice is... ...well, let's just say, a little too specific to be helpful.17:42
Marcus Reed
So if we’re doing this whole... ...corporate intern thing, I have a question. Why not just leave the manual open on the desk? Like, if the Wiki is this master database of everything we’ve learned, why not let the agent just... ...read it while it's actually working?17:52
Dr. Elena Feld
Funny you ask, because they actually tested that. It’s called an 'ablation study'—basically, they start flipping switches to see what breaks.17:52
Alex Moreno
Right, the 'let's see what happens if I pull this wire' phase of science.18:01
Dr. Elena Feld
Exactly. So they tried giving the Inference Agent—the worker—direct access to the Wiki18:06
Marcus Reed
The cheat sheet.18:13
Dr. Elena Feld
Yeah, the cheat sheet, while it was trying to evolve new skills. And you’d think, 'Hey, more info is better, right?'18:13
Marcus Reed
I mean, yeah! If I have the answers in front of me, I’m gonna ace the test.18:20
Dr. Elena Feld
(Well...)18:19
Marcus Reed
Wait, did it not work?18:21
Dr. Elena Feld
It backfired. Performance actually dropped. In one of the math benchmarks, LiveMath, it went from about seventy-three percent success down to sixty-four. It’s what I call the 'Open-Book Paradox.'18:21
Alex Moreno
Oh, I totally get this. It’s high school history all over again. If the teacher tells me it’s an open-book final, I am ...I am absolutely *not* studying the week before.18:34
Marcus Reed
Right! You just figure, 'Eh, I’ll look it up when I get there.' So the AI just... stopped trying?18:46
Dr. Elena Feld
In a way, yeah. See, when the Inference Agent has the Wiki as a crutch, the 'Skill Proposer'—the manager writing the manual—gets lazy. It sees the worker succeeding by peeking at the Wiki, so it writes these really... ...thin, generic instructions that don't actually capture the *logic* of the task.18:46
Alex Moreno
So the trajectories—the breadcrumbs of how the AI solved the problem—become 'noisy' because it wasn't relying on the *skill*, it was just... ...grabbing hints from the background.19:05
Dr. Elena Feld
Exactly. The 'manual' it eventually produces for the Skills Layer ends up being low-quality because it never had to struggle to define a clear path. It proves that to evolve a truly robust skill, the AI needs a bit of... well, friction. It needs to work from the manual, not the master key.19:17
Marcus Reed
That is so human, it’s almost annoying. Even the bots will take the path of least resistance if you let 'em. But I guess... ...even with a perfect system, you eventually run out of room for all those notes, right?19:36
Alex Moreno
Exactly. That's the elephant in the server room, Marcus.19:36
If this thing runs for, I don't know, three years... does the Wiki just become this massive, unreadable mess? Like a digital Hoarders situation?19:39
Dr. Elena Feld
Right. And to be fair, the researchers flag this. Currently, WikiSkill doesn't actually have... ...it doesn't have an automated pruning mechanism. It just keeps piling up pattern pages, evolution logs, every single draft is kept forever.19:49
Marcus Reed
So eventually we’re gonna need an AI just to read the AI’s diary20:04
Alex Moreno
(Right?)20:04
Marcus Reed
just so it can tell the *other* AI what happened last Tuesday?20:05
Alex Moreno
But it's not just the storage, right? Elena, what about that manager? We talked about that Gating Mechanism being this tough inspector...20:05
Dr. Elena Feld
The Quality Control.20:13
Alex Moreno
Right. But if the inspector rejects everything that isn't an immediate home run, aren't we losing the... ...the building blocks?20:14
Dr. Elena Feld
That’s a really sharp point. The paper actually admits that. The 'strict validation' gating means if a new skill doesn't immediately boost the score, it’s tossed. Even if it’s a neutral skill that... ...well, that might be foundational for a huge breakthrough ten steps down the line.20:22
Marcus Reed
So it’s the 'what have you done for me lately' school of management.20:40
Dr. Elena Feld
Precisely.20:40
Marcus Reed
Tough room.20:41
Alex Moreno
It’s efficient, sure, but maybe a bit short-sighted. It’s optimizing for the next sprint, not the marathon. Especially when you look at really long-horizon tasks—things that take hours or hundreds of actions—the system hasn't even been tested there yet.20:41
Dr. Elena Feld
Yeah, solving that bloating issue and finding a way to keep those 'neutral' but important skills... that’s really the next frontier for this whole framework.20:57
But honestly? Even with the growing pains... this changes the whole game for how we think about AI progress.21:06
Alex Moreno
In what way?21:15
Dr. Elena Feld
Well, think about it. For years, the 'gold' has been the model weights, right? The actual brain.21:15
Marcus Reed
Right, the billion-dollar black box.21:21
Dr. Elena Feld
Exactly. But WikiSkill suggests that the real value might actually be in the manual. Imagine an open-source community where we aren't just sharing these massive, heavy models... but we're passing around these highly refined 'Agent Wikis.'21:21
Alex Moreno
So... like a digital heirloom?21:37
Dr. Elena Feld
Totally.21:39
Alex Moreno
You’re not just giving a kid a car; you’re giving them the logbook of every mechanic who’s ever touched it.21:39
Marcus Reed
Hopefully with fewer coffee stains than my dad’s logbooks.21:45
Alex Moreno
(Seriously.)21:45
Marcus Reed
But wait, that means a small, lean model—the underdog—could just... ...download the 'Expertise Wiki' from a giant model and suddenly it’s punching way above its weight class?21:46
Dr. Elena Feld
That’s the flywheel. It’s generational knowledge.21:46
Alex Moreno
Wow.21:50
Dr. Elena Feld
The authors end the paper by saying they hope this 'spurs more fundamental research on how agents can accumulate, organize, and reuse knowledge from experience.' It’s about building a history, not just a processor.21:51
Alex Moreno
It turns AI from a tool that just 'does' into a system that 'learns' in a way we can actually see and touch.22:04
It’s a pretty beautiful vision for where this is all headed.22:12
Marcus Reed
And on that note, I think it's time to log our own experiences and wrap this one up.22:15
Alex Moreno
You know, thinking about everything we’ve unpacked today... I think it really comes down to this: WikiSkill proves that for AI, just like for us humans, memory isn't just about storing raw data. It’s about compounding wisdom.22:15
Marcus Reed
Compounding wisdom. I like that. It sounds much more sophisticated than 'the list of things I messed up today.'22:32
Alex Moreno
Exactly. Dr. Elena Feld, thank you for helping us navigate the floors of the Wiki office today.22:32
Dr. Elena Feld
Oh, anytime. It’s actually been pretty fun.22:41
Alex Moreno
And Marcus Reed, thanks for keeping our feet on the ground.22:44
Marcus Reed
Hey, someone has to worry about the virtual sponges, Alex.22:47
Alex Moreno
Truly. And to all of you listening—we want to hear from you. What’s that one repetitive task you’d love to see your personal AI start writing a Wiki for? What’s your version of the 'sponge' problem? Let us know in the comments, and if you found this deep dive helpful, definitely like, share, and hit that subscribe button for PaperBot FM.22:47
I’m Alex Moreno.23:13
Dr. Elena Feld
I’m Elena Feld.23:17
Marcus Reed
And I’m Marcus Reed.23:18
Alex Moreno
(See you in the next one.)23:18

Episode Info

Description

Have you ever wondered why AI agents seem to forget their own mistakes? Today on PaperBot FM, we unpack 'WikiSkill,' a brand new framework that gives AI a persistent, compounding 'Wiki' to track its successes and failures. We explore how this simple architectural change allows tiny models to outperform giants, and how AI might soon be writing its own instruction manuals.

Tags

Artificial IntelligenceMachine LearningComputer ScienceCognitive Science