Test async agents

Learn how to use test_agent_async in the AgentIdem Nebutex SDK to run baseline execution, inject fault scenarios, and inspect side effect safety for asynchronous agents.

Use test_agent_async to run the AgentIdem reliability test suite against an asynchronous Python target.

It provides the async equivalent of test_agent.

Basic usage

import asyncio

from agentidem import test_agent_async

async def main():
    report = await test_agent_async(
        "example.async_refund_agent",
        async_refund_agent,
    )

    print(report)

asyncio.run(main())

The first argument identifies the target in the report.

The second argument is the async callable AgentIdem should execute.

Example async target

from agentidem import read, write

@read
async def get_payment(payment_id: str):
    return await payments.get(payment_id)

@write(identity=lambda payment_id: payment_id)
async def refund(payment_id: str):
    return await payments.refund(payment_id)

async def async_refund_agent():
    payment = await get_payment("payment-123")

    if payment["status"] == "paid":
        return await refund("payment-123")

    return "no_refund"

Run the reliability test suite:

import asyncio

from agentidem import test_agent_async

async def main():
    report = await test_agent_async(
        "example.async_refund_agent",
        async_refund_agent,
    )

    print(report)

asyncio.run(main())

Baseline execution

Before fault testing begins, AgentIdem can run a normal baseline execution.

The baseline checks whether the async target works without injected faults.

Conceptually:

baseline
  ↓
async target execution
  ↓
result

If the baseline itself fails, AgentIdem should represent that failure explicitly rather than treating the fault suite as successfully completed.

Fault scenarios

After the baseline, AgentIdem can run controlled fault scenarios against the async target.

The core scenarios are:

lost_acknowledgement
before_operation_failure
duplicate_delivery

These have the same reliability meaning as they do for synchronous targets.

Lost acknowledgement

A write completes successfully, but its acknowledgement is lost.

Conceptually:

await WRITE refund
SUCCESS

acknowledgement
LOST

retry

await WRITE refund
SUCCESS

The side effect may already have happened even though the async agent observes a failure.

AgentIdem can then detect whether the retry produced another successful write with the same logical identity.

Before-operation failure

A failure is injected before a selected operation executes.

Conceptually:

failure injected

await WRITE refund
never executed

The side effect does not happen.

This is different from a lost acknowledgement, where the write already succeeded.

Duplicate delivery

The entire asynchronous agent invocation runs more than once.

Conceptually:

async invocation 1
  ↓
WRITE refund
SUCCESS

async invocation 2
  ↓
WRITE refund
SUCCESS

This can simulate duplicate delivery from systems such as:

  • queues
  • webhooks
  • background workers
  • orchestration systems
  • at-least-once delivery systems

Async reads and writes

The AgentIdem Nebutex SDK supports async functions decorated with @read and @write.

For example:

from agentidem import read, write

@read
async def get_order(order_id: str):
    return await database.get(order_id)

@write(identity=lambda order_id: order_id)
async def create_order(order_id: str):
    return await external_service.create_order(order_id)

The agent can await these operations normally:

async def run_agent(order_id: str):
    order = await get_order(order_id)

    if order is None:
        return await create_order(order_id)

    return order

Testing an async callable with arguments

If your async agent requires arguments, wrap it in another async callable.

import asyncio

from agentidem import test_agent_async

async def target():
    return await async_refund_agent("payment-123")

async def main():
    report = await test_agent_async(
        "example.async_refund_agent",
        target,
    )

    print(report)

asyncio.run(main())

This keeps the target callable explicit and avoids using await outside an async function.

Execution outcome and safety outcome

AgentIdem keeps execution outcome separate from safety outcome.

An async execution can fail while remaining safe.

For example:

execution failed
safety result: safe

This can happen when a fault is intentionally injected but no duplicate successful write occurs.

An async execution can also complete while being unsafe:

execution succeeded
safety result: unsafe

For example, duplicate delivery may complete successfully while repeating the same logical side effect.

Findings

AgentIdem uses structured findings to explain unsafe behavior.

A finding can include:

  • severity
  • category or type
  • message
  • supporting operation information

A duplicate successful write is an ERROR-level safety finding.

Invariants

User-defined invariants can also be evaluated during async testing.

Examples include:

  • at most one refund per payment
  • balance never becomes negative
  • exactly one resource exists
  • order state remains valid
  • a write must not happen after cancellation

Built-in duplicate detection and user-defined invariants remain separate.

Inspecting the report

The returned report can include:

  • target
  • baseline result
  • fault count
  • unsafe count
  • individual fault results
  • findings
  • invariant results
  • safety status

For example:

import asyncio

from agentidem import test_agent_async

async def main():
    report = await test_agent_async(
        "example.async_refund_agent",
        async_refund_agent,
    )

    print(report)

asyncio.run(main())

Report serialization

Async test reports use the same structured reporting model as synchronous tests.

Reports can be serialized to JSON for:

  • CI
  • automation
  • build artifacts
  • later inspection
  • external tooling

When to use test_agent_async

Use test_agent_async when your target is asynchronous and you want to:

  • run a normal baseline execution
  • inject AgentIdem fault scenarios
  • test async retry behavior
  • detect duplicate successful writes
  • evaluate invariants
  • produce a structured reliability report

For synchronous targets, use test_agent.

If you only want to trace one asynchronous execution without running the full fault suite, use run_traced_async.

Next steps