Error handling — bRRAIn Docs
Classify Platform SDK errors, react to each HTTP status correctly, set deadlines and retry only operations that are safe to repeat.
Error handling
Every SDK method returns error. What is inside falls into one of four kinds:
| Kind | How to recognize it | Example message |
| --- | --- | --- |
| APIError | errors.As(err, &apiErr) with apiErr *platformsdk.APIError | platformsdk: GET /internal/vault/file?path=... returned 404: {...} |
| Transport | errors.As(err, &urlErr) with urlErr *url.Error | ... context deadline exceeded (Client.Timeout exceeded while awaiting headers) |
| Client-side validation | A plain error returned before any request | platformsdk: Commit requires org_id, platformsdk: Grep requires pattern |
| Encode/decode | A plain error mentioning marshal or decode | platformsdk: decode GET /internal/... response: ... |
APIError carries Method, Path, StatusCode and the full response Body; its Error() string truncates the body to 200 characters. Never string-match err.Error() for status codes.
A classifier
type Class int
const (
Fatal Class = iota // fix the request or the configuration; do not retry
NotFound // absent resource; often a normal branch
Retryable // transient; retry with backoff if the call is safe to repeat
)
func Classify(err error) Class {
if errors.Is(err, context.Canceled) {
return Fatal // the caller gave up
}
var apiErr *platformsdk.APIError
if errors.As(err, &apiErr) {
switch apiErr.StatusCode {
case http.StatusNotFound:
return NotFound
case http.StatusTooManyRequests, http.StatusBadGateway,
http.StatusServiceUnavailable, http.StatusGatewayTimeout:
return Retryable
default: // 400, 401, 403, 409, 413, 416, 501 ...
return Fatal
}
}
var urlErr *url.Error
if errors.As(err, &urlErr) {
return Retryable // transport: dial, reset, client timeout
}
return Fatal // client-side validation or JSON decode: retrying cannot help
}
Status codes you will meet
| Status | Where it comes from | Reaction |
| --- | --- | --- |
| 400 | Missing fields, invalid JSON, a tool rejecting its arguments | Fix the request |
| 401 | Missing or wrong internal token | Configuration fault; alert, do not retry |
| 403 | Namespace guard, or the ingestion security gate | A policy answer; stop and report it |
| 404 | Missing file, folder, catalog entry or Operator ENV record | Usually a normal branch |
| 409 | A move refused by platform policy | Not allowed as a move |
| 413 | Ingestion size limit | A policy answer; split, roll over, or ask the administrator |
| 416 | ReadRange out of bounds | Fix the bounds |
| 501 | A backend this pod does not wire (for example Control().RoleOf on current pods) | Feature unavailable on this pod; fail closed |
| 502 | A backend failed: vault write, model or Handler, MCP tool (the tool's own status is inside the body) | Read the body; retry only if transient and safe |
502 is overloaded: it can mean "the Handler is still loading" (transient) or "this path is invalid" (permanent). Log the body. 501 is permanent for that pod version.
Deadlines and safe retries
Give every call a context deadline sized for the operation, and remember that Options.HTTPClient.Timeout caps every request regardless of the context. Retry only operations that are safe to repeat: reads, Stat, List, search, and whole-file overwrites. Do not blindly retry appends, ingestion commits, audit appends, notifications or model calls with side effects; a timeout does not tell you whether the first attempt landed.
// withRetry retries op on Retryable errors with exponential backoff.
// Use it only for operations that are safe to repeat (reads, overwrites).
func withRetry(ctx context.Context, attempts int, op func(context.Context) error) error {
delay := 250 * time.Millisecond
var err error
for i := 0; i < attempts; i++ {
if err = op(ctx); err == nil || Classify(err) != Retryable {
return err
}
select {
case <-ctx.Done():
return ctx.Err()
case <-time.After(delay):
}
delay *= 2
}
return err
}
func ReadWithRetry(ctx context.Context, c *platformsdk.Client, path string) ([]byte, error) {
var body []byte
err := withRetry(ctx, 4, func(ctx context.Context) error {
callCtx, cancel := context.WithTimeout(ctx, 10*time.Second)
defer cancel()
var err error
body, err = c.Vault().ReadBytes(callCtx, path)
return err
})
return body, err
}
Errors in your own API
Map SDK errors to responses for your interface deliberately:
NotFoundon something the user asked for: 404 with a clear message.- 403 and 413 policy answers: pass them on with the platform's reason in plain language.
- Configuration faults (401, 501): 500 or 503, with a log line holding the full
APIErrorand a message telling an administrator what to check. Retryableafter retries are exhausted: 503 with "try again shortly".
Never pass a raw APIError.Body to a browser without checking what it contains.
Testing error paths
Serve the pod's wire format from httptest and make it return the statuses you need to handle; the Quickstart shows a fake pod and a 401 test.
Next
- Examples
- bRRAIn Certified SDK Developer course: Module 7 covers errors, deadlines, retries and testing in depth.