Upcoming meetup

EBT #4: Practical AI for Founders, Developers & Operators

Date Jun 26, 2026 at 6:00 PM Place Lafayette, California

East Bay Tech brings founders, developers, and operators together for practical AI conversations grounded in real work.

Topics

These are early discussion starters for June 26, 2026. The list is building, and we will keep adding context before the meetup.

Current model snapshot: Artificial Analysis

  • Compared with meetup #3's May 22 snapshot, the frontier has shifted up and right: Claude Fable 5 and Claude Opus 4.8 now sit at the high-intelligence, high-cost edge, while GPT-5.5, Gemini 3.5 Flash, GLM-5.2, and Kimi K2.6 cluster close behind.
  • The most attractive quadrant is more crowded and more competitive than last month, with DeepSeek V4 Flash and Pro, MiMo-V2.5-Pro, and MiniMax-M3 showing how much pressure cheaper models are putting on the frontier labs.
  • Good room question: are teams actually changing their model choices as the cost-performance curve improves, or are product defaults and vendor trust still more important than raw benchmark economics?
Artificial Analysis chart from June 26, 2026 comparing intelligence versus cost to run across AI models.
Artificial Analysis model snapshot, captured June 26, 2026. The comparison point for meetup #4 is how quickly the attractive quadrant and frontier cluster changed since meetup #3.

Introducing Claude Opus 4.8

  • Anthropic says Opus 4.8 improves on Opus 4.7 across coding, agentic tasks, reasoning, and practical knowledge work while keeping regular pricing the same.
  • The launch is as much about workflow as model quality: Claude Code gets dynamic workflows for larger codebase-scale tasks, claude.ai gets effort controls, and Opus fast mode is cheaper than before.
  • Good room question: when models get better at long-running work, tool use, and self-checking, what should teams delegate versus keep under direct human control?
Anthropic benchmark chart comparing Claude Opus 4.8 with prior Claude models and other AI models across coding, agentic, reasoning, and knowledge-work tasks.
Anthropic's Opus 4.8 article includes a benchmark comparison across coding, agentic tasks, reasoning, and professional work.

GLM-5.2: Built for Long-Horizon Tasks

  • Z.ai positions GLM-5.2 as an MIT-licensed open model built for long-horizon work, with a solid 1M-token context and effort controls for coding tasks.
  • The launch is another sign that open models are competing less as cheap chatbots and more as serious agentic engineering systems, with claims around long-context coding, post-training, and local or hosted deployment.
  • Good room question: if open models can handle longer coding-agent trajectories, what parts of a team's AI stack should stay with closed frontier models, and what should move to open infrastructure?
Z.ai benchmark chart comparing GLM-5.2 with other frontier and open models on long-horizon coding tasks.
Z.ai's GLM-5.2 benchmarks frame open models as contenders for long-running coding-agent work, not just cheaper chat completions.

Previewing GPT-5.6 Sol: a next-generation model

  • OpenAI is beginning GPT-5.6 with a limited preview of Sol, Terra, and Luna, with broader availability planned after testing with trusted partners and coordination with the US government.
  • The announcement makes the Fable/Mythos debate more concrete: stronger cyber and biology capabilities are coming with heavier release controls, layered safeguards, and more automated red-teaming.
  • OpenAI also says GPT-5.6 Sol sets a new state of the art on Terminal-Bench 2.1, a coding benchmark for command-line workflows that require planning, iteration, and tool coordination.
  • Good room question: if frontier models are improving specifically at terminal-based agent work, how much of software engineering shifts from writing code to supervising long-running command-line workflows?
TerminalBench 2.1 score chart showing GPT-5.6 Sol Ultra at 91.9 percent, GPT-5.6 Sol at 88.8 percent, Claude Mythos 5 at 88.0 percent, GPT-5.6 Terra and Claude Fable 5 at 84.3 percent, GPT-5.5 at 83.4 percent, GPT-5.6 Luna at 82.5 percent, Claude Opus 4.8 at 78.9 percent, and Gemini 3.1 Pro Preview at 70.7 percent.
OpenAI's Sol announcement frames TerminalBench 2.1 as evidence that the GPT-5.6 family is getting stronger at terminal-based agent work.

Statement on the US government directive to suspend access to Fable 5 and Mythos 5

  • Anthropic says a US government export control directive required it to suspend Fable 5 and Mythos 5 access for all customers, while leaving access to other Anthropic models unaffected.
  • The post frames the dispute around model safeguards, jailbreak risk, and whether a narrow cybersecurity concern should justify recalling a widely deployed commercial model.
  • Good room question: what should a clear, fair process for pausing frontier model access look like when national security, customer reliability, and public technical evidence all collide?

Satya Nadella on tokenmaxxing and choosing the right model

  • Nadella points to a practical AI cost question inside serious organizations: not every task needs the most expensive or capable frontier model.
  • The useful shift is from "use more AI" to "match the model, tool, and workflow to the value of the work."
  • Good room question: how should teams decide when a frontier model is worth it, and when a cheaper, faster, or narrower system is the better engineering choice?

Matthew Berman on X

  • Matthew Berman highlights AWS CEO Matt Garman pushing back on the idea that companies can replace all junior developers with AI.
  • Contrast that with this YouTube Short as the sharper automation argument: if AI can handle more entry-level work, what still justifies hiring and training juniors?
  • It is a useful tension because junior roles are also how teams build taste, judgment, codebase context, and future senior engineers.
  • Good room question: where is Matt Garman right that companies still need junior developers, and where does AI genuinely change the entry-level career ladder?
YouTube Short thumbnail showing Ken Griffin on stage with text about AI agents and task completion.
The YouTube Short gives the automation side of the junior-developer debate a concrete visual prompt.

Come To

  • Meet smart people in the East Bay.
  • Share real use cases around practical AI.
  • Explore partnerships, projects, and business opportunities.
  • Connect with founders, developers, and operators across industries.

Format

  • Practical, thoughtful conversations for strong local connections.
  • Signal over hype.
  • Not a pitch night.
  • No aggressive selling or constant self-promotion.