A tool runner coordinates work that can fail in several places: input validation, permission checks, execution, artifact storage, and result delivery. A message saying “job failed” compresses all of those possibilities into one unhelpful sentence. Structured logs give each event named fields so a developer can follow a run, compare attempts, and understand where the intended result stopped being achievable.
The purpose is to answer operational questions with evidence. What was requested? Which stage ran? Was an artifact created? Can the operation be repeated safely? This guide develops an event design for a hypothetical document-processing runner. Use the structured logs workflow alongside it when connecting observability to a larger tool system.
Start with the questions an operator must answer
Write a short investigation scenario before adding log statements. Imagine that a document-processing task reports a failure after extracting its records. The operator needs to know whether extraction completed, whether the output was stored, and whether the caller received a response. Logging every line of execution would create more material without necessarily answering those questions.
Identify a small set of meaningful lifecycle events. For this runner, a useful starting set is request accepted, validation rejected, execution started, execution completed, artifact stored, and run finished. Add an event only when it explains a state transition or a decision that matters to someone operating the workflow.
Give the terminal event an explicit outcome. A run that was rejected before execution should differ from one that executed and failed. A canceled run should differ from one that exhausted its deadline. These distinctions support both incident investigation and a more accurate account of routine operation.
Use a shared record model where it helps
The OpenTelemetry Logs Data Model defines fields for event and observed timestamps, trace context, severity, body, resource, and attributes. It distinguishes the time an event occurred from the time a collection system observed it. It also separates information about the source of a log from attributes that vary for each event. The official OpenTelemetry Logs Data Model describes those meanings.
Use that reference to guide interoperability, then choose application fields that answer your team's questions. Do not label an arbitrary JSON object as an OpenTelemetry record merely because it contains a timestamp. If you export through an observability library, map your application concepts into that library's documented model deliberately.
Keep a small field dictionary in the project. Define each field's type, meaning, and allowed values. If one component uses attempt for a retry count and another uses it for a task identifier, a shared name creates confusion instead of consistency. Resolve those differences before building dashboards around them.
Separate the run from its individual attempts
Assign one identifier to the logical run and another to each execution attempt. A retry belongs to the same intended task but represents a distinct effort to complete it. Keep the identifiers stable as the work crosses internal boundaries. Include a parent identifier when a run creates child tasks that need to be followed independently.
Track resumed work
For the document example, a run might complete extraction on its first attempt but fail to store the output. A later attempt could resume storage if the architecture supports it. The logs should identify which stages were repeated and which artifacts were reused. Otherwise, an operator may assume that every retry repeated the entire workflow.
Preserve request context safely
Also preserve a correlation path to the original request without logging sensitive request contents. An opaque request identifier is often sufficient for joining authorized records. Decide who can resolve that identifier to source data and keep that access separate from ordinary log viewing.
Keep event names stable and messages readable
Choose event names that describe something that happened, such as tool.execution.completed. Keep the variable details in fields: tool name, stage, outcome, attempt identifier, and elapsed duration. A human-readable message can summarize the event, but automated analysis should not depend on extracting information from that sentence.
In the hypothetical runner, an execution-completed event could state the tool, run identifier, attempt identifier, output record count, and result artifact identifier. Each field has a purpose. The count helps explain an empty result; the artifact identifier helps locate the output; the attempt identifier connects the event to the correct execution.
Avoid embedding uncontrolled values in event names. A separate event name for every document title makes events difficult to group. Keep the name consistent and place approved contextual values in bounded fields. Review field additions as part of the interface because other tools may eventually depend on them.
Make errors useful without exposing the input
An error record should identify the failed stage, a stable error category, and an appropriate next action. For example, a validation failure can identify the field requiring correction. A storage failure can identify the storage operation and whether an output remains available locally. Avoid treating a long exception string as the complete error contract.
Log what is needed to investigate, not every available value. Exclude credentials, authorization headers, private document bodies, and other sensitive payloads by design. Review exception messages too, since a dependency may include input fragments or destinations in its text. Apply approved redaction before records leave the process boundary.
Set a policy for large diagnostic details. A short summary can remain in the routine log while restricted evidence is stored separately with an identifier. This keeps ordinary investigations readable and allows more detailed material to have appropriate access and retention controls.
Describe retries and deadlines explicitly
When an operation will retry, record the reason, the next attempt number, and the relevant remaining budget. When it will stop, record whether it reached an attempt limit, a deadline, or a non-retryable condition. The operator should not have to calculate the decision from scattered timestamps.
Choose severity according to a documented team policy. A temporary failure that is handled successfully may deserve different treatment from the final failure of the requested task. Keep the distinction consistent so an alert does not fire merely because a normal recovery path occurred. Preserve the underlying event even when it does not require an immediate alert.
Measure durations for stages that help explain latency. Queueing, execution, and artifact storage may deserve separate fields when the team needs to distinguish them. Label units explicitly and measure elapsed work using an appropriate runtime timer. Avoid mixing seconds and milliseconds under one field name.
Keep logs separate from the result contract
The caller should receive a documented result or error object even when detailed logs exist elsewhere. Do not require it to search logs to discover the output artifact or determine whether the operation succeeded. Logs explain the run; the result contract tells the caller what to do next.
For command-line tools, decide which output channel carries machine-readable results and which carries diagnostics. Ensure routine progress messages cannot corrupt the structured result stream. For a service, return the relevant run identifier with the response so an authorized operator can connect the caller's experience to internal evidence.
Design log-transport failure behavior as well. Decide how much buffering is acceptable, when records may be dropped, and how to expose a collection problem. A failure in the logging path should have an intentional effect on the tool's operation, chosen according to the importance of the evidence being recorded.
Keep routine queries simple
Draft the queries you expect to use before adding more fields. Can you find all events for one run, locate its terminal outcome, and group failures by tool and stage? Can you distinguish a retry that recovered from a task that finally stopped? These questions expose inconsistent names and missing identifiers early. Keep individual request identifiers available for investigation, while choosing bounded categories for routine summaries. Review stored volume against the value of each event so verbose development diagnostics do not become permanent operational noise by accident.
Verify the investigation path
Run a few meaningful scenarios: valid work, invalid input, a temporary failure, and an interruption after execution. For each, ask a teammate to reconstruct the outcome from the records. Check whether they can identify the final status, the relevant attempt, and any artifact that needs handling. Improve the record design where that reconstruction becomes guesswork.
Structured logs are valuable when they make the next operational decision easier. Use stable events, explicit identifiers, bounded contextual fields, and errors that identify a reasonable response. Connect them to the tool runner lifecycle and the tools API contract so execution, results, and evidence describe the same work.



