Robot Capability Profile Template for Arm RCF

Published

Use this template to describe what a robotic system can do in a defined use case, what evidence supports the claim, and what constraints govern operation. It is based on Arm’s Robotics Capability Framework Issue 1.0, but it is not an Arm assessment, certification, conformance test, or industry-standard rating.

The output is a capability profile, not one headline RL number. A system can show contextual perception, adaptive physical action, deliberative interaction, and reactive safety control at the same time. Averaging those observations into “an RL2.5 robot” would remove the distinctions the framework is intended to preserve.

If the system boundary itself is unclear, start with What Is Physical AI? and define where sensing, models, decisions, physical action, and human control sit before assessing capability.

Read Arm’s RL0-RL5 framework explained before using the template. The guide defines the six levels and their boundaries; this page supplies the working record.

How to use this profile

  1. Define one use case and operating context. Create separate profiles when the task, environment, supervision, software, or hardware differs materially.
  2. Select only relevant capabilities. NOT APPLICABLE needs a reason; it should not hide missing evidence.
  3. Record observed behavior and a source that can be reviewed.
  4. Compare the evidence with Arm’s level-specific expectations. Do not infer a level from a product name, model label, or demonstration video alone.
  5. Record limitations, human authority, fallback, safety, and security for each capability.
  6. Date the evidence. A software, model, sensor, or environment update can invalidate the profile.

A. System and use-case record

Copy this table into an engineering document, procurement record, issue tracker, or spreadsheet.

FieldEntry
System name and version
Robot type and embodiment
Assessed use case
Deployment environment
Mission or task boundary
Allowed operating conditions
Excluded conditions
Human role and supervision
Intervention and override path
Hardware, sensors, and effectors
Software, models, and policies
On-robot, edge, and cloud dependencies
Network and disconnection assumptions
Applicable safety/security process
Assessor and evidence owners
Assessment date and next review

The physical AI technology stack can help identify the hardware, runtime, simulation, model, policy, and deployment layers that belong in the record.

B. Robotics-specific capability profile

The capability names below come from Arm RCF Issue 1.0. The “observed alignment” field is deliberately capability-specific. Use INSUFFICIENT EVIDENCE when a claim cannot be reproduced or bounded.

Arm capabilityRequired behavior in this use caseObserved behaviorPossible RL alignmentEvidenceConfidenceLimitations and failure casesHuman/fallback control
PerceptionINSUFFICIENT EVIDENCE
Localization & MappingINSUFFICIENT EVIDENCE
Physical ActionINSUFFICIENT EVIDENCE
Decision-MakingINSUFFICIENT EVIDENCE
Learning & AdaptationINSUFFICIENT EVIDENCE
InteractionINSUFFICIENT EVIDENCE
Safety & Trust GovernanceINSUFFICIENT EVIDENCE

What counts as evidence?

Useful evidence can include:

  • a repeatable test protocol and result;
  • logs showing inputs, decisions, actions, uncertainty, and intervention;
  • a safety or assurance artifact scoped to the capability;
  • a versioned model card, software release, interface contract, or configuration;
  • a failure-injection, recovery, or rollback result;
  • a vendor document that states an exact capability and product status;
  • an independently reproducible demonstration with defined conditions.

A marketing phrase such as “cognitive,” “general-purpose,” or “self-learning” is a claim to investigate, not evidence of alignment.

C. General-compute foundation

Arm treats general-compute capabilities as the platform properties required to execute, isolate, observe, update, secure, and manage robotics workloads. They support the robotics profile but are not an intelligence level themselves.

General-compute consideration from Arm RCFRequirementImplementation/evidenceConstraint or gapOwner
Compute execution and workload management
Determinism and timing
Memory
Connectivity
Monitoring and diagnostics
Resilience and fault containment
Configuration and software lifecycle
Updateability and rollback
Orchestration
Safety mechanisms
Cybersecurity

Latency, compute placement, memory, power, determinism, and safety are important architecture constraints in Arm’s framework material. They should not be converted into independent RL scores. Record how they enable or constrain a capability.

D. Ecosystem-facing profile

Arm’s taxonomy names four ecosystem-facing capabilities. These describe whether the system can integrate, evolve, and scale without prescribing one vendor or architecture.

Arm capabilityRequirementEvidenceConstraint or dependencyOwner
Interoperability
System Integration
Stability
Scalability

Use the robot API and SDK directory to locate official interfaces, but do not treat an SDK listing as proof that two components interoperate in the assessed system.

E. Capability architecture: System 1, System 2, and System 3

Arm uses System 1, System 2, and System 3 as architectural roles, not mandatory chips or deployment locations:

LayerArm roleCapability or component in this systemDeployment locationLatency/determinism needBehavior if unavailable
System 1Reflex: fast, automatic, stimulus-driven; includes hard real-time control
System 2Tactical autonomy: mission-bounded reasoning and adaptation
System 3Strategic autonomy: learning, training, shared knowledge, and long-horizon improvement

On-robot, edge, and cloud are deployment choices. A System 2 capability can run on the robot or a local edge service if the latency, reliability, and mission boundary permit it. System 3 can be embedded or distributed. Do not label every cloud workload System 3 or every onboard controller System 1 without examining its role.

F. Safety, security, and assurance

Arm treats safety and security as cross-cutting properties and assurance expectations. Functional safety also depends on the general-compute foundation. This template does not replace a hazard analysis, safety case, cybersecurity assessment, or certification process.

CheckEvidence or decisionGap/constraintOwner
Operating envelope and prohibited actions defined
Uncertainty and degraded-mode behavior defined
Human authority, override, and emergency stop tested
Timing, watchdog, isolation, and safe-state behavior tested
Input spoofing, identity, authorization, and communications assessed
Model, data, update, and supply-chain integrity assessed
Decision provenance and audit evidence retained where required
Learned or updated behavior requires validation before deployment
Rollback and recovery exercised
Applicable standards and regulatory duties assessed separately

The sim-to-real workflow provides practical failure, sensor, and rollback testing. Simulation evidence must still be validated on the physical system within safe limits.

G. Capability conclusion

Do not produce a single overall score. Summarize instead:

Conclusion fieldEntry
Capabilities with strong evidence
Capabilities with partial evidence
Capabilities with insufficient evidence
Highest-risk assumptions
Required human supervision
Blocking platform or ecosystem dependencies
Tests required before broader deployment
Changes that trigger reassessment
Evidence review date

If stakeholders insist on an RL statement, keep it narrow: “For capability X in use case Y under conditions Z, the observed evidence appears consistent with selected expectations at RLn.” Label it non-certifying and assessment-specific. Do not turn it into a product-wide badge.

Example of a bounded observation

Weak statement:

Robot A is RL3.

Stronger statement:

In the documented warehouse-picking task, version 2.4 selected among predefined grasp behaviors using object state and short-horizon context. That evidence may be consistent with contextual decision-making expectations, while physical action remained bounded and recovery required operator intervention. No product-wide RL level was assigned.

The stronger version exposes the task, software version, capability, evidence boundary, and human role. It can be challenged and retested.

What this template deliberately does not do

  • certify conformance with Arm RCF;
  • imply Arm approval;
  • map products from vendor marketing alone;
  • replace ISO, IEC, IEEE, SAE, regulatory, safety, or cybersecurity work;
  • equate RL5 with AGI or uncontrolled self-modification;
  • prescribe Arm processors or a particular model, middleware, or deployment topology;
  • average capability observations into a single robot score.

The robot foundation model guide explains why model availability, open weights, policy execution, and commercial deployment are separate evidence states. A sophisticated model does not establish an integrated robot’s capability level.

Sources