Video summary

How AI Is Rewriting 40 Years of Audit Work | Aryo Patel & Tinah Hong, Andera

Main summary

Key takeaways

Technology

Tech/product focus: AI-native audit automation (EnderA)

  • EnderA is building an AI-native platform to automate audit and financial assurance, starting with SOC control testing (the transcript repeatedly says “socks,” likely referring to SOCs).
  • Core claim: their tech achieves “100% coverage” of SOC controls for the largest Fortune 500 companies, and they work with advanced internal audit teams to “reimagine” audit workflows.

Why SOC control testing (“the wedge”) matters

  • SOC/SOC-like data is positioned as unusually valuable because it lets auditors map a company’s entire financial ecosystem:
    • SOC testing supposedly checks that material reported items (e.g., revenue) are within about 1% accuracy.
    • Auditors are described as seeing 100% of the story, while accounting/IT/compliance see only parts.
  • The company argues this domain provides the best starting point for broader automation, with an eventual longer-term vision of an “AI audit brain” that makes judgment calls across the full financial ecosystem.

What’s different about their approach (engineering & architecture)

  • The main bottleneck is not “just better models,” but:
    • Search and exploration over distributed audit data
    • Context links not scaling enough for audit-grade reasoning
  • They emphasize a two-stage workflow:
    1. Interpret what’s happening in the data
    2. Answer specific SOC questions
  • To make interpretation efficient, they do heavy lifting “up front” so the AI requires minimal reasoning effort at query time.
  • Excel parsing is a major technical differentiator:
    • They built a custom Excel parsing engine to expose how spreadsheets work “under the hood.”
    • They also write more complex features back into Excel (automation that outperforms what they claim is available via open-source tooling).
  • They tailor agent execution:
    • Prefer an agent representation/harness that can work well “in distribution” with how the underlying models are trained.
    • They rebuilt agent harnesses multiple times because what’s optimal changes as model training/post-training changes.

Performance/experience claims vs competitors

  • They report customers see controls go from ~two weeks to a few hours for previously slow tasks.
  • They accept a trade-off:
    • slower processing in exchange for handling more complex problems and reducing the risk of nuanced misinterpretation.
  • They argue their generalization comes from data interpretation at scale, not hardcoded rules:
    • Competitors are said to hardcode rules per control set, which becomes unscalable.
    • Their approach is intended to be generalizable across other audit verticals (e.g., security, financial statements, lender due diligence).

Market adoption & trust (deal-making)

  • A key theme: trust is the gating factor for replacing Big Four work.
  • EnderA reportedly wins deals by:
    • building credibility with industry veterans
    • maintaining SLA adherence and operational continuity (customers’ testing process doesn’t “change” unexpectedly)
    • having a team split of ~50% auditors and 50% engineers
  • Proof point mentioned: some customers replaced Big Four co-sourcing teams and gave EnderA 100% of SOC testing.
  • They also mention customers creating case studies for EnderA proactively.

“Why now?” (what changed in the market)

  • Auditors are described as labor-constrained and already inclined to believe audit work “should get automated,” but historically lacked a clear path to implementation.
  • The transcript claims multiple forces aligned:
    • stronger enterprise pressure to adopt AI and cut costs
    • board/audit committee incentives to increase AI adoption
    • LLMs enabling one general approach to data expressed in many formats (reducing the need for many bespoke automations)
  • A “crossroads” point: earlier attempts with RPA or outsourcing didn’t work well due to data nuance, often leading to outsourcing rather than in-house automation.

2025/technology milestones and product scaling strategy

  • They describe 2025 emphasis as:
    • building platform tech to prove it scales
  • Technical bets (examples):
    • custom Excel parsing engine
    • intentional agent harness design focused on stable data representation and adaptability as models evolve
  • They note it took about a year and a half to reach market to build the underlying tech first.

How they distinguish from AI “labs” and future model progress

  • They argue model improvements from labs don’t automatically translate into better audit outcomes because:
    • the harder part is data search/exploration + audit-specific judgment/rationalization
    • “off-the-shelf” agent harnesses (e.g., generic coding agent setups) don’t match audit needs
      • e.g., memory/compaction design decisions require audit-specific harnessing
  • Their evaluation results (as claimed): when models improve, their evals only rise by ~1–2%, implying no major breakthroughs solely from lab model upgrades.

Big Four response & EnderA’s positioning

  • The Big Four collectively employ ~1.5M people annually and audit most public companies.
  • EnderA’s view: Big Four must adapt or cannibalize themselves if they simply build AI tools without changing value capture.
  • Observed adaptations:
    • partnering with model/lab providers (e.g., KPMG with Anthropic; EY spending on internal tools)
    • rethinking offshore labor and headcount in an AI-first workflow
    • offering additional services to offset freed capacity
    • involving startups to manage deployment transitions

Longer-term vision: what must be true

  • Two major “must haves”:
    1. Data interpretation at scale becomes “more crackable” (implying Excel-domain complexity is a key research gap vs more coding-like tasks where models are post-trained more aggressively).
    2. Trust must be built quickly—historically built over hundreds of years by Big Four—attempting to compress that timeline via senior industry partnerships.
  • Their near-to-mid roadmap includes expanding upstream from SOC testing into broader judgment tasks (timeline described as 3–4 years for “productionizing” and expanding use cases).

Main speakers/sources (as stated)

  • Bucky Moore (host)
  • Cat Zang (host)
  • Ario Patel (co-founder, EnderA)
  • Tina Hong (co-founder, EnderA)

Original video