Video summary

تازه‌ترین اخبار هوش مصنوعی: مدل‌های جدید مایکروسافت و انویدیا، و جی پی تی

Main summary

Key takeaways

News and Commentary

Summary of the video’s main points (AI news & model updates)

Opening / context

  • The presenter returns after a long gap and frames the episode as an AI news update.
  • They express condolences and hope for better days, referencing recent tragic events (including the “Mina school”).
  • They also note Iran’s internet was reconnected, which delayed earlier publishing.

Microsoft’s announcements (major focus)

7 new model releases

The presenter highlights 7 new Microsoft model releases, covering:

  • language/thinking models
  • coding
  • image generation
  • speech-to-text transcription
  • audio generation

Positioning strategy

  • The presenter argues Microsoft is aiming for reliable, capable, state-of-the-art “independent” performance, not necessarily the absolute strongest/newest model in every category—especially for coding.

Arena-style leaderboard notes

In an Arena-like comparison where users vote/compare quality:

  • Microsoft image model is ranked D
    • described as strong, “after GPT”.
  • Speech transcription model is emphasized as:
    • 2–3x faster
    • relatively low error
    • but the presenter claims it does not yet support Persian
      • based on shown language support (English and others mentioned).

Microsoft “Solara” platform + agentic devices

Solara: beyond chatbots

  • Microsoft introduces Solara, intended to enable agentic capabilities beyond chat—including real-world actions on devices.

New device concepts

  • A table-top conversational device (Alexa-like) that can perform task actions such as:
    • calendar operations
    • file-related actions
  • A badge-like device with a screen/camera that can scan/record conversations and images for workplace/industry use cases, such as:
    • doctor scanning barcodes
    • industrial scanning

Debate on hardware vs software

  • The presenter notes there’s debate whether these capabilities require a dedicated device.
  • They also say Microsoft is trialing market response.

NVIDIA + Microsoft collaboration: “RTX Spark”

  • NVIDIA and Microsoft describe a high-end Windows PC/supercomputer concept called RTX Spark.
  • Key technical claim:
    • CPU and GPU are unified on one chip
    • using Unified Memory
    • targeting up to 128GB, similar to Apple’s approach
  • Intended benefits:
    • stronger performance for AI inference
    • including image/video/game/editing workloads

Pricing: unknown

  • Speculation: ~$3,000–$4,000+

  • Microsoft also announces Surface Ultra using the same chip:

    • expected to be high performance
    • cost still unknown

Coding / leaderboard positioning (including a “Chinese market update”)

The presenter gives an overview of “best model” options for coding/app development:

  • Cloud Opus 4.7 / 4.7 Thinking:
    • presented as the top choice for coding + creativity/thinking
  • Mentions strong open-source coding, e.g.:
    • K / Alibaba’s Qwen 3.7
  • Mentions mixed rankings:
    • some Gemini/GPT variants are described as lower than Cloud in the presenter’s assessment
  • DeepSeek:
    • described as creating hype but viewed as not matching expectations
  • Gemini variants:
    • framed as good multi-purpose models
    • especially for image analysis to generate outputs
    • but less strong for coding than Cloud

More model releases & trends (open-source, smaller + faster)

  • Nemotron Strata (NVIDIA)

    • oriented toward long-term agent work rather than everyday chat
    • open-source
    • very large (~550B parameters), likely requiring cloud/API rather than local use
  • Gemma 4 (Google)

    • open-source
    • “unified architecture” (presented as replacing older “multimodal” framing with a single architecture)
    • emphasized as able to run on smaller hardware, with example needs like:
      • laptop class: ~16GB RAM
      • GPU class: ~16GB VRAM
    • logic performance emphasized; smaller footprint claimed
  • MiniMax / Ems

    • described as open-source coding models
    • said to match strong proprietary options on many benchmarks
  • Composer inside Cursor IDE

    • Cursor is framed as an agent-like coding environment with codebase awareness
    • Composer 2.5 is described as:
      • comparable coding accuracy
      • lower cost
      • especially versus Cursor subscription tiers that can spike under heavy usage

Image generation: new approach and key models

Structure-based generation (vs diffusion-style)

The presenter compares:

  • diffusion-style image generation
  • a newer structure-based method:
    • models convert prompts into a structured concept hierarchy
    • decide spatial placement of objects
    • enable precise edits (change one element without disturbing everything else)

Model contenders

  • GPT-Image / Google Nano-like models (referenced as contenders)
  • A newer set of models described as:
    • good quality
    • improved control
    • including strong consistent character generation behavior

Ideogram

  • Ideogram
    • emphasized as open-source
    • strong prompt following
    • claimed to edit photos well (add/remove elements)
    • referenced as doing well in Arena comparisons (strong, though not necessarily #1)

Video generation and editing

  • Focus on a Google video model that:

    • generates videos from prompts
    • can edit existing video
      • e.g., change gender/design
      • reposition objects/people
    • described as learning “world/physics-like” structure rather than only producing visually plausible output
  • The presenter highlights an example where the model generates video aligned to:

    • provided graphics
    • written instructions —presented as a more “reasoned/structured” workflow than earlier approaches.

Sound / voice models

  • Mentions an audio/voice model (from “Meez” and “Lebs” in subtitles) that claims:
    • open-source availability
    • control over emotions and voice tone
    • English-only limitation currently
  • Notes sample code and reinforces the trend toward smaller open-source models that can run locally.

Google Gemini Flash + agentic work + app building

  • Gemini 3.5 Flash

    • described as 2–3x faster than heavier models
    • keeps similar accuracy for certain tasks
    • useful for “identity work” / folder-based labeling examples
      • e.g., naming photos based on content + task
  • Google AI Studio

    • described as enabling Android app building inside the browser:
      • prompts generate code + preview
      • later connect to a phone for installation/testing
  • Similar iOS app creation concept mentioned via related tooling.

Cloud coding / “Codex”-style tooling

  • Explains Cloud Code / Codex-type products:
    • users don’t need deep coding knowledge
    • the tool breaks tasks into steps and builds apps end-to-end
  • Compared with earlier coding chat assistants that required more manual fixing:
    • this approach aims to reduce “getting stuck”
  • Mentions plugin/skill ecosystems and an example iOS workflow created through these systems.

Personal app anecdote

  • The presenter says they built an iOS app called “Dreamweaver” using these tools:
    • dream journaling
    • dream interpretation inspired by Carl Jung psychology
  • Claims AI coding accelerated:
    • frontend/UI development
    • backend interpretation logic partly handled by the presenter

Google’s inbox-focused AI + additional research

  • Highlights an idea: AI inside the inbox that can:

    • read/summarize emails
    • propose actions
    • voice-read messages
    • enable replying
  • Additional Google research/tools mentioned:

    • “Notebook”-style summarization of sources (text/PDF/video) and generation of study assets (including podcasts/graphs)
    • CoScientist (agent for scientific research):
      • runs through literature
      • generates hypotheses
      • critiques/improves hypotheses
      • example claims include:
        • proposing a drug regimen for rare leukemia
        • findings related to liver fibrosis
        • and a vision-related problem

Presenters / contributors mentioned

  • Microsoft (model/platform announcements)
  • Google (announcements, Gemini models, AI Studio, research systems)
  • NVIDIA (RTX Spark collaboration; Nemotron Strata)
  • Open-source model developers/companies referenced
    • Alibaba (Qwen 3.7)
    • DeepSeek
    • MiniMax
    • Kimi
    • Ideogram
    • Cursor/Composer ecosystem
    • Entropy (Codex/Cloud Code described)
  • Carl Jung (referenced via “Dreamweaver” interpretation framework)
  • The video’s presenter (speaker): not named in the subtitles (addressed as “friends” / “I am back…” only)

Original video