TL;DR: Naming a concept is easy. Building it is hard. RAPSπΌRAPSRust CLI for Autodesk Platform Services.View in glossaryβs advocacy layer β raps-kernel β handles auth, rate limiting, circuit breaking, failure classification, and bulk orchestration for 16 Autodesk APIs. Here is what it took.
The Gap Between Concept and Code
In part 1, we established why vendor APIs are hostile by design. In part 2, we introduced the advocacy layer as a pattern β an intermediary that absorbs APIπAPIInterface for software components to communicate.View in glossary complexity so application code does not have to.
This post is different. This is the implementation post. Five real engineering problems, five solutions, and the architectural choices that made them possible in Rustπ¦RustSystems programming language known for safety.View in glossary.
Architecture: The Crate Dependency Graph
RAPS is not a monolith. It is a workspace of focused crates, each owning one API domain, all sharing a common kernel.
Every domain crate depends on raps-kernel for auth, HTTP, rate limiting, and error handling. No domain crate implements its own tokenποΈTokenCredential for API authentication.View in glossary refresh. No domain crate manages its own retry logic. The kernel owns all of it.
This is the structural payoff: write the hard parts once, get them right once, and every API domain inherits the solution.
Problem 1: OAuth Race Conditions
The scenario
You have 10 concurrent API calls running. The access token expires. All 10 calls detect the expiry and attempt to refresh simultaneously. The first refresh succeeds. The second refresh invalidates the first token. Now 8 of your 10 calls fail with 401s, even though you just refreshed.
This is not hypothetical. This is what happens in production with naive token refresh.
What the official SDKs do
Nothing. They treat token refresh as a per-request concern. If two requests refresh at the same time, that is your problem.
The RAPS solution
The kernel uses a Mutex + tokio::sync::Notify pattern to ensure exactly one refresh happens at a time:
// Conceptual pattern (simplified)://// 1. Caller checks token expiry// 2. If expired, acquires refresh lock// 3. First caller to acquire lock performs the refresh// 4. All other callers wait on Notify// 5. First caller stores new token, notifies all waiters// 6. Waiters read the fresh token and proceed//// Result: N concurrent callers, exactly 1 refresh,// 0 token invalidation races.
The key insight: token refresh is not a per-request operation. It is a shared resource that must be synchronized. The type system enforces this β you cannot access the token without going through the guard.
Result: Zero token invalidation under concurrent load. Every bulk operation that fires 50+ concurrent requests relies on this.
Problem 2: Chunked Uploads
The scenario
An architect needs to upload a 2GB Revitπ RevitAutodesk's BIM software for architecture and construction.View in glossary model. The official SDKπ§°SDKLibrary for building applications on a platform.View in glossary does a single POST. At 1.8GB, the connection drops. The entire upload fails. Start over.
Adaptive chunked uploads with parallel workers and resumable state:
raps object upload my-bucket large-model.rvt# Automatically:# - Detects file size, selects chunk strategy# - Splits into optimal chunks# - Uploads 5 chunks concurrently# - Retries failed chunks (3 attempts per chunk)# - Resumes from last successful chunk on network drop
The implementation breaks down into four layers:
Chunked Upload Pipeline
1
Size Detection
Files under 100MB: single PUT. Over 100MB: chunked upload initiated. Chunk size adapts to file size and available bandwidth.
2
Session Init
Creates a resumable upload session with OSS. Returns a session URI that survives connection drops.
3
Parallel Upload
5 concurrent workers, each uploading one chunk. Failed chunks retry 3 times with exponential backoff. Workers pull from a shared chunk queue.
4
Finalization
Commits the upload session. Verifies integrity via content hash. Reports total time and throughput.
The critical detail: resumable state is persisted to disk. If the process crashes at chunk 47 of 100, restarting the commandβΆοΈCommandInstruction executed by a CLI tool.View in glossary picks up at chunk 48. No wasted bandwidth, no re-uploading gigabytes of data.
Problem 3: Failure Classification
The scenario
An API returns HTTP 403. What does that mean?
Your token lacks the required scope? Re-auth and retry.
Your account does not have permission on this projectπProjectContainer for folders and files within a hub.View in glossary? Skip this project, continue with the restπRESTWeb service architecture style using HTTP.View in glossary.
The API is misconfigured? Log and alert.
Rate limited with a 403 instead of 429? (Yes, some APIs do this.) Back off and retry.
A single HTTP status code maps to multiple failure modes that require completely different handling strategies. If you treat them all the same, you either retry everything (wasting time on permission errors) or fail everything (stopping bulk operations on the first hiccup).
The RAPS solution
The kernel classifies every API failure into one of 13 types, each with its own recovery strategy:
Failure Classification Matrix
Auth Expired
Refresh token, retry once
Rate Limited
Exponential backoff with jitter
Server Error
Circuit breaker β fail fast after threshold
Permission Denied
Skip, log, continue (bulk ops)
Network Error
Retry with timeout escalation
Not Found
No retry β resource does not exist
Validation Error
No retry β caller error
Each failure type maps to a backoff strategy:
No backoff β auth refresh, immediate single retry
Exponential with jitter β rate limits, transient server errors
Circuit breaker β sustained server errors trip the breaker, all subsequent calls to that endpoint fail fast for a cooldown period
The circuit breaker is especially important for bulk operations. If an API endpoint is down, you do not want 1,700 calls hammering it. You want the first 5 failures to trip the breaker, then every subsequent call fails immediately with a clear message: βendpoint circuit breaker open, retry after cooldown.β
Problem 4: Bulk Operations at Scale
The scenario
A company acquires another company. 1,700 projects. Every project needs the same 50 users added with the correct roles. The Autodesk web UI lets you add users one project at a time.
1,700 projects x 50 users = 85,000 clicks. That is not an exaggeration. That is the math.
The RAPS solution
raps admin user add "$ACCOUNT" "user@company.com" --role project_admin# 1,700 projects. 10 concurrent workers. Rate-limit aware. Resumable.# What takes 2+ hours of clicking: 45 seconds.
The bulk operation engine in raps-admin orchestrates this with four properties:
Concurrency Control
10 concurrent workers by default. Each worker processes one project at a time. The worker count respects API rate limits β the kernel throttles automatically when quota drops below 10%.
Resumable State
Progress is checkpointed to disk. If the operation is interrupted at project 847 of 1,700, resuming starts at project 848. No duplicate work, no missed projects.
Failure Isolation
A permission error on project 312 does not stop the remaining 1,388 projects. Failed projects are logged and reported at the end. The operation continues.
Progress Tracking
Real-time progress bar with ETA. Shows completed, failed, and remaining counts. Operation ID for status queries on long-running jobs.
This is where every prior engineering investment pays off. Bulk operations do not implement their own auth refresh β the kernel handles it. They do not implement their own rate limiting β the kernel handles it. They do not implement their own retry logic β the failure classifier handles it.
The bulk engine is ~800 lines of code. Without the kernel, it would be 5,000+.
Problem 5: Permission Clone and Export
The scenario
You set up a project with a complex permission structure β 30 folders, each with specific access levels for different roles. Now you need to replicate that structure across 50 new projects. The ACC web UI offers no βclone permissionsβ feature.
The RAPS solution
# Clone permission structure from one project to anotherraps admin folder set-permissions "$ACCOUNT" \ --project "$TARGET_PROJECT" \ --clone-from "$SOURCE_PROJECT"# Export permissions to CSV for auditraps admin folder set-permissions "$ACCOUNT" \ --project "$PROJECT" \ --export permissions.csv
This feature composes three capabilities:
Read the complete permission tree of the source project (folder hierarchy + per-folder ACLs)
Map the folder structure between source and target (folders may have different IDs but matching paths)
Apply the permission set to the target, with conflict resolution and dry-run mode
The export-to-CSVπCSVTabular data format for spreadsheets.View in glossary path exists for compliance teams who need audit trails. Every permission change is logged with timestamp, user, action, and before/after state.
Why Rust
This is not a language advocacy post. But the choice of Rust is load-bearing for this project, and here is why.
Memory safety
Bulk operations run for minutes, processing thousands of API responses. In a GC language, memory pressure causes pauses. In Rust, memory is freed deterministically. No GC pauses during a 1,700-project user addition.
async/await
tokio gives us concurrent API calls without thread-per-request overhead. 10 concurrent workers sharing a single runtime, with backpressure handled by the channel.
Type system
API response parsing is validated at compile time. When Autodesk changes a field from string to object (and they do), the compiler catches it before users do. 579 commits in 2 months at this velocity because the type system prevents regression.
Single binary
No runtime dependencies. No Python version conflicts. No Node.js compatibility matrix. One binary, every platform. Users run curl -sSL install.rapscli.xyz | sh and they are done.
The Reuse Payoff
This is the core insight of the advocacy layer pattern, and it bears repeating:
Write the kernel once. Deploy it everywhere.
Same auth logic in:
β CLI (interactive use)
β Cloud SaaS (multi-tenant API gateway)
β Python bindings (data science scripts)
β MCP server (AI assistant integration)
Same rate limiting in:
β Single-user CLI sessions
β 50-concurrent-request bulk ops
β Multi-tenant cloud with per-tenant quotas
βCI/CDπCI/CDAutomated build, test, and deployment pipelines.View in glossary pipelines running unattended
No duplication. No drift. One fix in the kernel fixes every consumer.
When we fixed a race condition in token refresh, it was fixed in the CLI, the cloud service, the Python bindings, and the MCP server. One commit. When we tuned the circuit breaker thresholds, every consumer got the improvement. When we added a new failure classification for a previously-unseen Autodesk error response, every API domain inherited the handling automatically.
This is what makes the advocacy layer an architecture, not just a library. It is not a helper you call β it is the substrate everything runs on.
What Comes Next
The advocacy layer enables everything else. It absorbs the complexity of 16 APIs so that higher-level features β bulk operations, permission management, model translationβοΈTranslation JobBackground process converting CAD files to viewable formats.View in glossary, reality capture β can focus on domain logic instead of fighting HTTP.
Next in the series: the specific problems this architecture solves for real organizations β starting with βBIM 360π΅BIM 360Legacy Autodesk construction platform (predecessor to ACC).View in glossary to ACC: The Migration Nobody Planned For.β