Square
Following
Hot
News
Profile

jolestar

vip
Active for: 8.7y
Peak Tier 5
No content yet
0
Following
7
Followers
40
Liked
DeepSeek doesn’t seem cheap either. My primary models are gpt 20x and glm 20x, with DeepSeek-Flash mainly serving as a fallback. Could it be that these two are often overloaded, so DeepSeek ends up silently shouldering the load? 😅
GLM-2.49%
I saw Ayman Nadeem’s article “Plan Mode Is Dead.” But personally, I still prefer having the Agent generate a planning document for me to review first—it doesn’t necessarily have to be Plan mode.
Because there are multiple ways to solve the same problem, I need to review its choice of approach, how it implements feature switching, how it handles compatibility, and how it plans to roll things out step by step.
I remember one time when, because I didn’t carefully review the proposed approach, the Agent came up with an extremely complicated progressive feature-flag rollout strategy. The flags were
post-image
MODE+3.92%
Holon(@holonrun) originated from several real pain points I encountered in my daily use of AI Agents. It evolved from a small toy built for personal use into the primary system I use every day. But it still has a long way to go before becoming a mature general-purpose tool.
If you’re interested in this kind of persistent, asynchronous, work-item-driven Agent, I’d be very happy for you to try it and share your feedback.
Adding a Small Decision Model to the Agent Scheduling System
I tested Jev and picked up the semantic scheduling system I had wanted to build for @holonrun some time ago.
At the time, I was refactoring Holon's scheduling system with the goal of allowing Agents to switch between multiple work items like humans do: when one work item needs to wait for an external process, switch to another work item and continue execution.
But I ran into several challenges. After an Agent receives an input, it is difficult for the system to determine using static rules alone:
* Should this GitHub event be inserte
The Agent wrote a script and started it in the background. It forgot to increment the cursor in the while loop, creating an infinite loop that filled up the disk with output and then crashed itself 😅. This kind of thing also seems difficult to prevent in the harness.
One side effect of the incident in which Guo Degang was fined for altering a revolutionary song is that many music services are removing revolutionary songs. Who can guarantee that their version hasn’t been altered? 😅
During summer vacation, I installed a coding agent on both kids’ computers and watched how they used it.
My older son has experience with Scratch programming, has ideas of his own, and has made several creative Scratch games. After trying a few times, he gave up on having AI generate applications directly. His main reasons were that what AI generated wasn’t what he wanted and that he couldn’t clearly explain the ideas in his head to AI. He also felt that what AI made wasn’t really his own work, so he couldn’t experience the feeling of creation.
In the end, he had AI modify his Scratch files. S
After using AI a lot, I’ve found that I don’t really feel like talking much anymore, even online. I haven’t posted a tweet in ages.
Back when I programmed the old-fashioned way, writing too much code would make me feel pent-up, so I’d write articles or chat in groups. But after using AI, writing code is basically chatting, and it seems I’ve used up my output quota—I no longer feel the urge to produce anything.
I’ve also lost patience for chatting with people. After just a couple of exchanges, I want to say, “Why don’t you go chat with AI instead?” or “I think I’ll go chat with AI.”
But after a
The model often makes mistakes when using `rg`, with an error rate of about 10%. The issue is that `rg`'s handling of `-rn` is inconsistent with `grep`, and the model is more familiar with `grep`, so it frequently uses it incorrectly. In the AI era, new tools replacing old ones should seamlessly accept all inputs from the old tools, especially those that LLMs are familiar with.
post-image
GPT's subscription was inexplicably canceled by Google Play, and now it seems there is no subscription entry on the Android version of GPT? Has anyone encountered a similar issue?
holon v0.19.0 has been released
This version includes an integrated web UI, and also completely reconstructed the underlying storage system, the API, and the event system, filling many AI-driven gaps.
Originally, it used JSONL, but as the data volume grew, it became difficult to maintain.
Coupled with AI programming habits of grepping and adding when not found, resulting in many redundant APIs and events.
So everything was moved into an SQLite database, but unexpectedly, as data accumulated, the database exceeded 4GB, leading to performance issues.
Therefore, further optimization was
Codex device token login suddenly requires phone number verification? And I found that the OpenAI phone number can't be found in account settings? It seems there's no place to change it.
What Basic Toolset Does an Agent Need
Seeing everyone discussing the question of Agent toolsets—does providing a shell solve everything? After working on holon, I realize it's not that simple.
Read: Why abandon Read/Glob, switch entirely to shell
holon’s toolset has gone through several versions, ultimately discarding specialized tools like Read (file reading) and Glob (pattern search) provided by Claude Code, relying entirely on shell for reading and searching. This aligns with Codex’s approach—Codex’s ExecCommand is straightforward: reading files with cat, searching code with rg, witho
Assign a milestone to the Codex plan, then keep adding issues into it, and it will keep working. Unfortunately, my pace of adding issues can't keep up with its implementation speed 😅
After communicating with GPT more and more, I’ve gotten used to using the word “wrap up” too. Once some tasks are done, but there are still a few leftover bits and pieces that haven’t been handled, telling it to wrap up the remaining things feels very natural. I even forgot how I used to say it before I started using the word “wrap up” 😅.
How do these kinds of words get inserted? Although GPT 5.5 is already powerful enough, issues like this always make people doubt its reliability😅
Codex Plus's weekly limit is approaching, having multiple windows open for a long time without closing caused iTerm to consume dozens of GB of memory, and the disk was also filled up by the Agent's worktree, constantly prompting to clean up the window. So I was forced to restart my computer, opened a Codex to use the remaining quota to clean up the disk, planning to take a break and rest. As a result, I found that Codex reset the limit!!😅
post-image
In the AI Coding era, good programming habits still matter
Recently, I was working on an Agent benchmark and found that you can't simply evaluate the complexity of a programming task for AI from a developer's perspective.
For example, a refactoring task: splitting a large file of several thousand lines into more than ten small modules based on functionality.
This task isn't really difficult for a developer; the main work involves moving code, organizing imports, and verifying compilation, which even beginners can handle.
So I thought of using a simple task for benchmarking, but the res