Video summary
Why Is My AI So Slow? Go Faster with EvoX AI Harness
Main summary
Key takeaways
Product reviewed: EvoX AI Harness
A “coding and productivity assistant” desktop-style app (shown running on Mac in the video) that uses cloud AI models plus optional local workflows. It’s positioned as faster than local generation, with tools for:
- Coding agents
- Self-evolution / learning from mistakes
- Swarm-style autonomous task execution
Key features mentioned
Coding + productivity assistant (multiple work modes)
The app includes multiple work modes/sections:
- Chat
- Co-work (more integrated tasks)
- Dedicated coding section
Cloud execution with a free credit system
- Starts with $15 worth of credits
- Creator reports 1,500 credits available at signup
Model choice (including “thinking modes”)
Models mentioned include:
- DeepSeek V4 Flash
- GPT Sol
- Luna
- Sol
- Terror
- Grok 4.6 (SpaceX)
- Kimi K3
Thinking effort modes:
- extra high
- high
- minimal
- off
The video mainly compares low and extra high, noting that low can reduce quality.
Cloud vs local workflow
The assistant can operate in two ways:
- Local system mode: generate files on your machine
- Remote execution mode: generate/operate remotely “into a computer”
It can also:
- Create and organize project folders
- Write generated files into those folders
Safety/automation controls for code changes
The app includes controls around file writing and edits, such as:
- Prompting for approval to write files
- Options like:
- auto-accept edits
- confirm steps
- run smart accept edits
- “Similar commands” option to reduce repeated confirmations
Beta features
- Beta swarm agent (autonomous background development/testing)
- One-click migration of projects
Skills system
The assistant can load specialized skills depending on the task (e.g., interactive visual demos/games). The creator claims it can:
- save/reuse memory/skills to improve future outputs
Self-evolution / self-improvement
The video shows a “self-evolution” progress indicator, e.g.:
“94% away from self-evolution”
This implies the system tracks progress toward evolving/improving its future capability and skills.
Reasoning visibility differs by model
The creator states:
- With closed-weight models, reasoning is hidden
- With Kimi K3, some reasoning can be seen
What the reviewer tested (and outcomes)
1) Speed comparison: local vs cloud
- Local Kimi K3 was described as extremely slow:
- Example: 7,000 tokens at ~6.62k/sec
- Cloud EvoX using Kimi K3 was “a lot faster” for a similar task
2) “Solar system” HTML generation (Kimi K3 vs GPT Sol)
- Kimi K3: produced a working, good-looking interactive solar system
- GPT Sol:
- Initially produced a blank screen
- Then it fixed the code after adjusting imports and logic
Conclusion: Kimi K3 looked better visually, while GPT Sol required fixes (but recovered successfully).
3) Interactive game-like task: spaceship + lasers + explosions concept
The creator attempted a more advanced interactive feature.
- GPT Sol:
- Created a flyable spaceship and working laser shooting
- Laser visuals had a direction issue (described as going the “wrong direction” / vertically)
- Still playable and had no runtime errors
- Kimi K3:
- Lasers/planets appeared, but the spaceship wasn’t properly visible/controllable
Conclusion: GPT Sol “won” for gameplay/control responsiveness.
4) Extra-high benchmark: photo-realistic real-time face (cloud vs local)
- Cloud (GPT Sol, extra high):
- Finished in about 1 minute
- Iterated after defects (e.g., missing mouth)
- Later corrected to include mouth/lips with improved rendering/shaders
- Local comparison (GLM 5.3 quantized):
- At low thinking: 77,000 tokens and ~2 hours
- Also showed a WebGL2 requirement bug message (though rendering may have still happened)
Explicit conclusion: EvoX/cloud was dramatically faster (1 minute vs 2 hours) with visible iteration to fix defects.
Swarm agent test (autonomous coding/testing)
- Swarm agent enabled with Ultra cool set to auto
- Prompted it to generate an implementation for a new model: DeepSseek V4.1
- The swarm reportedly:
- Checked out existing implementations/skills
- Compared against prior models’ code (e.g., older DeepSeek V4 and Quen 4)
- Requested permission to clone MLX source
- Ran a background plan including:
- writing/reading files
- creating tests
- comparing diffs
- producing implementation files locally
- Result: creator reports implementation files were successfully generated
- UX detail: notifications keep the user informed so they can “relax and watch YouTube.”
Pros (as stated/shown)
- Much faster than local generation
- Example cited: face render ~1 minute cloud vs ~2 hours local
- Strong code-writing workflow
- Creates files, updates code, and corrects errors (e.g., blank-screen fix, face-mouth fix)
- Useful autonomy via swarm
- Background planning, testing, cloning dependencies, generating implementations
- Self-improvement concept (“self-evolution”) with a visible progression indicator
- Model flexibility
- Switch between models and thinking modes; reasoning visibility varies by model
- Skills system
- Loads task-appropriate capabilities (e.g., game/visual skills)
Cons / negatives mentioned
- Closed-weight model reasoning is hidden
- Quality can drop at low thinking effort
- Example: GPT Sol at low produced a blank screen first
- Local execution can be extremely slow
- Example cited: ~2 hours run
- Some output artifacts/bugs can appear
- Laser direction visual artifact (lasers moving vertically)
- Face render initial missing mouth + occasional shader/render quirks
- Local GLM showed a WebGL2 requirement message
Comparisons made
EvoX cloud vs local generation speed
Multiple tasks showed drastic time savings—especially the face render (1 min vs 2 hrs).
Model comparisons within EvoX
Kimi K3 vs GPT Sol:
- Solar system: Kimi K3 looked better; GPT Sol needed fixes
- Spaceship/lasers: GPT Sol performed better (flyable spaceship)
- Face render at extra high: GPT Sol produced a strong result quickly
Reasoning visibility
- Kimi K3: more reasoning visible
- Closed-weight models: reasoning hidden
Unique points list (distinct claims/features/outcomes mentioned)
- EvoX is a latest coding & productivity assistant.
- Designed to learn from mistakes and be self-evolving.
- Provides free API credits ($15, reported 1,500 credits).
- Supports multiple models including DeepSeek V4 Flash, GPT Sol, Kimi K3, Grok 4.6, etc.
- Works across channels: communicates via Slack and Telegram (claimed).
- Includes beta swarm and one-click project migration.
- Desktop app UX includes Chat, Co-work, and Dedicated coding sections.
- Supports working with local files and remote computer access.
- Creates project folders and writes outputs to them.
- Supports model choice including thinking modes (extra high/high/minimal/off).
- Local Kimi K3 performance was extremely slow (example: 7,000 tokens, 6.62k/sec).
- Kimi K3 cloud run is “much faster.”
- GPT Sol may hide reasoning; Kimi K3 can show reasoning.
- GPT Sol at low thinking initially produced blank screen, then fixed code.
- Kimi K3 produced a working solar system; GPT Sol needed import/logic correction.
- GPT Sol can fix issues after code edits (responsive iteration).
- Spaceship/lasers test: - GPT Sol produced flyable spaceship and lasers - Kimi K3 attempt lacked visible/controllable spaceship
- GPT Sol had a laser direction artifact but was playable with no runtime errors.
- “Extra high” face render with GPT Sol completed in about 1 minute.
- Local face render using GLM 5.3 quantized took 77,000 tokens and ~2 hours.
- Local GLM had a WebGL2 bug message (though rendering possibly happened).
- Face render initially missing mouth; user prompt led to correction.
- Self-evolution indicator shown: “94% away from self-evolution.”
- Self-evolution implies it saves/uses a “gene” and improves next runs.
- Swarm agent test involved generating DeepSeek V4.1 implementation.
- Swarm autonomously did planning, cloned deps (MLX), checked/used other models (DeepSeek V4, Quen 4), and generated tests/implementation.
- Swarm produced implementation files in the user’s folder.
- Swarm sends notifications so user doesn’t need to actively watch.
- Overall verdict in the video: GPT Sol “wins” several higher-quality comparisons; extra-high cloud face generation is best.
Speaker-specific views (attributed)
- Single main speaker/reviewer (video narrator):
- Provides comparisons, benchmarks, and qualitative judgments.
- Claims about reasoning visibility, speed differences, and self-evolution come from their demonstrations.
- No other distinct speakers are clearly present in the subtitles.