Error handling — bRRAIn Docs

Classify Platform SDK errors, react to each HTTP status correctly, set deadlines and retry only operations that are safe to repeat.

Error handling

Every SDK method returns error. What is inside falls into one of four kinds:

| Kind | How to recognize it | Example message | | --- | --- | --- | | APIError | errors.As(err, &apiErr) with apiErr *platformsdk.APIError | platformsdk: GET /internal/vault/file?path=... returned 404: {...} | | Transport | errors.As(err, &urlErr) with urlErr *url.Error | ... context deadline exceeded (Client.Timeout exceeded while awaiting headers) | | Client-side validation | A plain error returned before any request | platformsdk: Commit requires org_id, platformsdk: Grep requires pattern | | Encode/decode | A plain error mentioning marshal or decode | platformsdk: decode GET /internal/... response: ... |

APIError carries Method, Path, StatusCode and the full response Body; its Error() string truncates the body to 200 characters. Never string-match err.Error() for status codes.

A classifier

type Class int

const (
	Fatal     Class = iota // fix the request or the configuration; do not retry
	NotFound               // absent resource; often a normal branch
	Retryable              // transient; retry with backoff if the call is safe to repeat
)

func Classify(err error) Class {
	if errors.Is(err, context.Canceled) {
		return Fatal // the caller gave up
	}
	var apiErr *platformsdk.APIError
	if errors.As(err, &apiErr) {
		switch apiErr.StatusCode {
		case http.StatusNotFound:
			return NotFound
		case http.StatusTooManyRequests, http.StatusBadGateway,
			http.StatusServiceUnavailable, http.StatusGatewayTimeout:
			return Retryable
		default: // 400, 401, 403, 409, 413, 416, 501 ...
			return Fatal
		}
	}
	var urlErr *url.Error
	if errors.As(err, &urlErr) {
		return Retryable // transport: dial, reset, client timeout
	}
	return Fatal // client-side validation or JSON decode: retrying cannot help
}

Status codes you will meet

| Status | Where it comes from | Reaction | | --- | --- | --- | | 400 | Missing fields, invalid JSON, a tool rejecting its arguments | Fix the request | | 401 | Missing or wrong internal token | Configuration fault; alert, do not retry | | 403 | Namespace guard, or the ingestion security gate | A policy answer; stop and report it | | 404 | Missing file, folder, catalog entry or Operator ENV record | Usually a normal branch | | 409 | A move refused by platform policy | Not allowed as a move | | 413 | Ingestion size limit | A policy answer; split, roll over, or ask the administrator | | 416 | ReadRange out of bounds | Fix the bounds | | 501 | A backend this pod does not wire (for example Control().RoleOf on current pods) | Feature unavailable on this pod; fail closed | | 502 | A backend failed: vault write, model or Handler, MCP tool (the tool's own status is inside the body) | Read the body; retry only if transient and safe |

502 is overloaded: it can mean "the Handler is still loading" (transient) or "this path is invalid" (permanent). Log the body. 501 is permanent for that pod version.

Deadlines and safe retries

Give every call a context deadline sized for the operation, and remember that Options.HTTPClient.Timeout caps every request regardless of the context. Retry only operations that are safe to repeat: reads, Stat, List, search, and whole-file overwrites. Do not blindly retry appends, ingestion commits, audit appends, notifications or model calls with side effects; a timeout does not tell you whether the first attempt landed.

// withRetry retries op on Retryable errors with exponential backoff.
// Use it only for operations that are safe to repeat (reads, overwrites).
func withRetry(ctx context.Context, attempts int, op func(context.Context) error) error {
	delay := 250 * time.Millisecond
	var err error
	for i := 0; i < attempts; i++ {
		if err = op(ctx); err == nil || Classify(err) != Retryable {
			return err
		}
		select {
		case <-ctx.Done():
			return ctx.Err()
		case <-time.After(delay):
		}
		delay *= 2
	}
	return err
}

func ReadWithRetry(ctx context.Context, c *platformsdk.Client, path string) ([]byte, error) {
	var body []byte
	err := withRetry(ctx, 4, func(ctx context.Context) error {
		callCtx, cancel := context.WithTimeout(ctx, 10*time.Second)
		defer cancel()
		var err error
		body, err = c.Vault().ReadBytes(callCtx, path)
		return err
	})
	return body, err
}

Errors in your own API

Map SDK errors to responses for your interface deliberately:

  • NotFound on something the user asked for: 404 with a clear message.
  • 403 and 413 policy answers: pass them on with the platform's reason in plain language.
  • Configuration faults (401, 501): 500 or 503, with a log line holding the full APIError and a message telling an administrator what to check.
  • Retryable after retries are exhausted: 503 with "try again shortly".

Never pass a raw APIError.Body to a browser without checking what it contains.

Testing error paths

Serve the pod's wire format from httptest and make it return the statuses you need to handle; the Quickstart shows a fake pod and a 401 test.

Next