To Data & Beyond

To Data & Beyond The place for continuous learning on Advanced AI, GenAI, and Data Science

We are putting the final touches on Prompt Engineering 2.0, the second edition of my Hands-on Prompt Engineering book, w...
07/10/2026

We are putting the final touches on Prompt Engineering 2.0, the second edition of my Hands-on Prompt Engineering book, which was first published in 2024.

This new edition reflects how prompt engineering has changed over the last two years. It moves beyond prompting standalone LLMs and focuses much more on prompting and designing agentic systems.

It covers topics such as agent prompting patterns, Agent Skills, building and controlling agent loops, and designing prompts for more reliable multi-step workflows.

Stay tuned! Prompt Engineering 2.0 goes live tomorrow, and I’ll be offering a special launch discount for the first 100 copies.

06/10/2026

Yann LeCun's latest talk at ETH Zürich:

"You should not work on LLMs. At least if you're in academia, you should absolutely not work on LLMs. There is nothing you can bring to the table.

If you're interested in making real progress in AI, in grounded AI for the real world, or in physical AI, don't work on LLMs, and don't work on generative models either!

So, as you can probably guess, this does not make me very popular in Silicon Valley."

New live workshop: Claude Cowork MasterClass: From Prompts to Deliverables & Automated WorkflowsClaude Cowork is the tab...
05/10/2026

New live workshop: Claude Cowork MasterClass: From Prompts to Deliverables & Automated Workflows

Claude Cowork is the tab in your Claude desktop app where Claude reads your folders, uses your connectors, runs scheduled tasks, and hands back finished work: a cleaned-up folder, a populated spreadsheet, a weekly digest.

Most people never make the jump from prompting to delegating. The productivity gains stay theoretical. This workshop closes that gap in 120 minutes. You'll leave with two real workflows running on your actual machine.

What we'll build together:

✅ Workflow #1 — The messy folder cleanup

✅ Workflow #2 — Your first recurring task

Date & Time: Sun 18 Oct 2026 19:00 - 21:00 EEST

Early bird price is 32$ till Sunday, 11 Oct; after that, it will be 40$.

Book your seat: https://www.tickettailor.com/events/todatabeyond/2405067

Harness engineering is the new core AI skill. This repo teaches it for free.14 lectures, 8 projects, 4 frontier harness ...
03/10/2026

Harness engineering is the new core AI skill. This repo teaches it for free.
14 lectures, 8 projects, 4 frontier harness breakdowns:

When an agent run goes wrong, the setup around the model is usually the weak point. Anthropic tested this. Same model, same prompt, two runs.

A solo agent run cost $9, took 20 minutes, and produced a game that didn't work. A run with a planner, generator, and evaluator cost $200, took 6 hours, and shipped a playable game.

This course teaches you to build that kind of setup yourself.

You start by running one task prompt-only, then rules-first, and see the gap yourself. From there you add session handoffs, scope limits, and self-verification. By project 8 you're wiring maker-checker loops into a graph with rollback and human approval.

Each project's solution becomes the next one's starter, so you watch the agent get more reliable as the harness grows.

The breakdowns also show how Claude Code, Codex, DeepSeek, and Pi do it in production.

Link in the comments!

New free live workshop on Building Decision Layers for AI Agents with Jev.A lot of agent workflows still use general-pur...
02/10/2026

New free live workshop on Building Decision Layers for AI Agents with Jev.

A lot of agent workflows still use general-purpose LLMs for small decisions that do not require generated text. Things like:
→ Which model should handle this request?
→ Is this tool call risky?
→ Is this retrieved result relevant?
→ Should the agent continue, stop, or ask for human approval?

In this workshop, Youssef Hosni will explain how Jev can be used as a decision layer around an AI agent, where the model handles bounded semantic judgments while normal code keeps control over the final action.

We’ll cover practical use cases including:
- model routing
- tool-risk gating
- retrieval and relevance checks
- verification and supervision
- escalation and human approval

Youssef Hosni will also walk through a hands-on implementation and share results from my own experiments comparing Jev with a general-purpose LLM on latency, cost, and decision behavior.

The goal is not only to understand how Jev works, but also where it actually fits inside an agent architecture, when deterministic code is still better, and when you still need a full reasoning model.

The workshop is completely free, but seats are limited.

Book your seat from the link in the comments.

MASSIVE reveal from Google! Its new flagship, Gemini 4 Argon, outscores GPT-6 Astra and Claude Opus 5.5 on most benchmar...
01/10/2026

MASSIVE reveal from Google!

Its new flagship, Gemini 4 Argon, outscores GPT-6 Astra and Claude Opus 5.5 on most benchmarks.

- It beats GPT-6 Astra and Claude Opus 5.5 on some super important industry benchmarks.

- Its widest lead in legal work, 19.6% on Harvey's Legal Agent Benchmark against 6.7% for Anthropic's Claude Fable 5.1.

- Output limit jumps from 64K to 1M tokens, an industry-leading ceiling,

- Only 3 groups have it today. The first is Google's own staff, vetted cyber defenders such as government agencies and security companies and trusted testers giving Google feedback.

- Inside Google, Argon agents freed over 300 TiB of data-center memory, with 500 TiB to 1 PiB of total savings estimated, and made a Rust port of the libgav1 video decoder 2.7x faster by replacing 32K lines of SIMD code.

30/09/2026

By 6 December, you will have built your own Claude agent system, one layer at a time, over six live Sundays. That is what we build together in Claude Agent Engineering, my new 6-week live cohort, which starts on Sunday 1 November.

Each week adds one layer on top of the last, so by the end the pieces work as a single system:

- Week 1: Claude Code, set up as your engineering environment

- Week 2: Claude Cowork, for longer multi-step work with clean handoffs

- Week 3: Agent Skills, turning a workflow into a Skill that triggers reliably

- Week 4: Claude Agent SDK, rebuilding it as an agent you control in code

- Week 5: Agent loops that run, verify, retry and stop

- Week 6: Context engineering (write, select, compress, isolate) for long-running work

Every Sunday is a 3-hour live session at 18:00 EET, and there is a 1-hour office hour every Tuesday, which comes to 24 live hours in total. You also get the recordings, slides, code and a private Slack group, and the cohort is capped at 20 seats.

The early-bird price is $320 instead of $400 with the code CLAUDECOURSE20, and it ends on Wednesday 7 October.

The enrollment link is in the comments.

28/09/2026

You can use jevgrep - a research agent CLI powered by jev from typesafeai that reduces your coding agent cost by 40% (verified on SWE-bench)

You can also use the built in skill so your coding agent knows to use jg for context collection. This works because coding agents typically spend 30-60% of all its tokens on research to collect context before writing a single line of code. The actual code generation tokens are tiny.

To install, just send claude/codex this exact repo: https://github.com/dzhng/jevgrep

A new hands-on article on Jev is now available on To Data & BeyondJev Clearly Explained: I Built a Decision Layer for an...
26/09/2026

A new hands-on article on Jev is now available on To Data & Beyond

Jev Clearly Explained: I Built a Decision Layer for an AI Agent

The question I wanted to answer was simple: where does a decision model like Jev actually fit inside an agent, and what do you gain by using it instead of sending every small judgment to a general-purpose LLM?

I built a tool-call risk gate where Jev evaluates risk, destructiveness, and whether human approval is needed, while Python keeps control over the final allow / review/block policy. Then I ran the same workflow through Jev 1.13.0 and GPT-5.6 Terra, with 40 measured calls per model.

In this experiment:

- Jev median latency: 242.6 ms

- Terra median latency: 1,511.5 ms

- Jev p95: 280.5 ms

- Terra p95: 2,039.6 ms

- Estimated list-price difference: about 82× in Jev’s favor

The numbers were useful, but the more interesting part was the behavior on ambiguous commands. A production Kubernetes delete, for example, did not produce a clean risk classification in Jev, yet the surrounding policy still returned review in all 10 runs.

That is the part we found most useful: a model can be uncertain while the system around it remains conservative and predictable.

The article includes the architecture, runnable code, real terminal results, repeated benchmarks, and the cases where I would still choose deterministic code or a general-purpose LLM.

Link in the comments!

Mozilla just published a 91-page report and it says open-weight AI is now only about 4.4 months behind the frontier.- 8 ...
17/09/2026

Mozilla just published a 91-page report and it says open-weight AI is now only about 4.4 months behind the frontier.

- 8 of OpenRouter's 10 most-used models by August token volume were open-weight, and 7 were Chinese-built, while DeepSeek became the first open model to lead the platform in weekly requests.

- However, the economics are almost upside down. Open models handled roughly 20% of measured OpenRouter usage but captured only about 4% of model-layer revenue in the cited 2025 window.

Mozilla attributes much of that mismatch to pricing, with closed models costing roughly 6x more per call at about 90% capability parity.

The comparison shows why high usage does not necessarily translate into high revenue when one class of models is dramatically cheaper.

Mozilla also warns that the revenue measurement is older than its 2026 usage data, so the current revenue split may be different.

- Open models are spreading faster than they reach production: 79% of surveyed developers use them, but only 51% of open-model deployments reach production versus 63% for closed models.

- DeepSeek showed that a new pretraining run may no longer be necessary for a major capability jump, gaining roughly 10 index points and then another 8 through post-training passes.

-Even model diversification may provide less protection than assumed: Kimi K3 and Claude Fable 5 had a 0.72 per-task failure correlation, meaning supposed backup models often fail on the same problems.

Osoite

Helsinki

Ilmoitukset

Saat tiedon ensimmäisenä: lähetämme sinulle sähköpostia, kun To Data & Beyond julkaisee uutisia ja tarjouksia. Sähköpostiosoitettasi ei käytetä muihin tarkoituksiin, ja voit perua tilauksen milloin tahansa.

Ota Yhteyttä: Koulu

Lähetä viesti: To Data & Beyond:

Oikopolut

Jaa

Kategoria