docs/spec-compiler-part-4-partitioned-data.md

← all documents
not measured0highlighted source passages0/0PBTs with runs0draft PRs
Not estimatedcombined confidence estimate0 considered · 0 missing
How combined confidence is computed

max(0, 1 − sum of PBT residual-risk estimates). Uses each distinct available PBT’s latest-run estimate and its own sampling scope; PBTs without estimates are reported and omitted. An observed violation remains the suite result; when other PBTs have estimates, their estimate is still shown. Assumes no independence. This is a combined point estimate, not a statistical confidence bound or deployment reliability. Uncovered passages are outside its scope.

No PBTs registered for this file.

Code coverageNot measured

Coverage added over the baseline, grouped by PBT. Only files with gains appear below.

No added coverage recorded.

Files without added coverage and unmeasured PBTs
Source fileBaseline coverageBaseline + inputContributing input
–

docs/spec-compiler-part-4-partitioned-data.md · pinned revision 48615bc5925ef4b9db8b4550b5d4322933cf4b7b

1

Loom Compiler Part 4: Logical Domains And Data Views

3

This document specifies the software-side logical-domain and derived-data-view contract used at Loom thread boundaries. The canonical ABI is intentionally small: a thread definition owns behavior and exactly one logical-domain kind, while each launch supplies that domain's root parameters and passes all data as ordinary typed operands. The two domain kinds are a zero-based dense Cartesian domain and a responsibility-tracked dynamic work domain.

10

The earlier thread_axis, staticGrid*, dataflow.map_info, dataflow.partition_domain, and dataflow.partition_layout design was removed because it duplicated domain, schedule, and boundary facts already owned by the Structured Program Candidate, launch ABI, and SystemMapping.

15

1. Authority And Scope

17

Part 4 owns:

19
  • the closed logical-domain kind of a thread definition;
  • the logical-coordinate interpretation of a thread definition's trailing index block arguments;
  • launch-domain cardinality rules;
  • dynamic work-item identity, publication, retirement, and termination;
  • source induction-variable reconstruction; and
  • the boundary between ordinary software data views and physical Mapping.
27

Part 4 does not own:

29
  • thread or SpatialCore outlining decisions;
  • parallel, temporal, tiling, interchange, vector, or unroll choices;
  • physical AccCore selection, route, Tag, reservation, or topology;
  • a second memory-transfer or partitioned-data ABI; or
  • a Runtime queue implementation, channel-lifecycle protocol, work-stealing policy, or multi-process atomic-memory protocol.
36

2. Logical Domain Kinds

38

Every dataflow.thread definition carries exactly one closed domain value:

40
ThreadDomain =
    DenseRectangular
  | DynamicWork { work_item_arg_ordinal }
46

The canonical attribute spellings are:

48
#dataflow.thread_domain<dense>
#dataflow.thread_domain<dynamic_work, work_item_arg = N>
53

This value is program semantics. It is neither Mapping policy nor a scheduler configuration. work_item_arg_ordinal identifies exactly one ordinary function_type input; it is absent for DenseRectangular. No string kind, extension registry, or second work-domain object exists.

58

2.1 Dense Rectangular Domain

60

A dataflow.thread entry block has this canonical shape:

62
(ordinary_args..., thread_ctrl : none, coord_0 : index, ..., coord_{K-1} : index)
66

The suffix length K is the coordinate rank and is derived from this block shape. No rank, axis-kind, grid, layout, or topology attribute duplicates it.

69

Each dataflow.thread.launch supplies exactly K index extents. The dynamic instance set is:

72
[0, extent_0) x ... x [0, extent_{K-1})
76

Every extent is non-negative. Rank zero creates exactly one instance. If any extent is zero, the domain is empty and the collective completion token retires after launch dependencies without executing a thread body. Static verification rejects a provably negative extent; runtime admission rejects a dynamic negative value before creating any instance.

82

The coordinate tuple identifies an instance but defines no row-major linear order, issue order, physical grid, or hardware topology. If program semantics require a linear id, the Structured Program Candidate computes it explicitly from coordinates and extents.

87

2.2 Dynamic Work Domain

89

A DynamicWork thread has no coordinate suffix. Its designated ordinary argument is the typed work-item payload; all other ordinary arguments are launch captures reused by every item in that domain instance. A caller-side dataflow.thread.launch supplies exactly one root payload and no extents.

94

Each dynamic item has the stable runtime identity:

96
WorkItemId =
  (domain_instance, root_or_parent_item_id, child_launch_ordinal)
101

The root uses the distinguished parent Root and ordinal zero. A child's ordinal is the zero-based program-order occurrence of dataflow.work.spawn within its parent item execution. Payload equality does not merge items, and queue position, worker identity, AccCore choice, address, and wall-clock time never enter identity.

107

dataflow.work.spawn %payload is legal only while executing a DynamicWork thread. Its operand type must equal the designated work-item argument type. It publishes one child to the current domain; it is not an arbitrary nested dataflow.thread.launch, creates no new domain, returns no handle, and cannot target another thread definition.

113

The first DynamicWork profile has no channel endpoints. A DynamicWork thread cannot create, send, receive, capture, or bind !dataflow.channel<T>, and a graph launch owned by such a thread cannot bind graph stream ports to a channel. Work-list payload, child publication, active responsibility, memory, and collective thread completion already form one complete model; adding a WorkItemId message correspondence or non-affine channel relation would be a second dynamic identity mechanism. A later profile may reopen the boundary only with a concrete program that cannot use a dense channel domain, explicit payload termination, or memory-backed work sharing.

123

The semantic termination authority is the domain's active responsibility set:

125
  • launch admission acquires the root responsibility before the root is visible;
  • the root source closes immediately after that one root publication, so later items can arise only through registered child spawn;
  • spawn atomically acquires a child responsibility before making that child visible to any worker;
  • a queued, in-flight, or executing item retains exactly one responsibility;
  • completion of dataflow.thread.yield retires the current item exactly once; and
  • the launch's collective !dataflow.thread_token retires exactly when the root source is closed and the active responsibility set is empty.
136

An active-count implementation is a derived cache of that set, not a second semantic authority. Publication-before-retirement prevents a transient zero from terminating a domain while a child is becoming visible. One logical coordinator remains responsible for this atomic transfer. The Runtime ABI may place this responsibility kernel behind a bounded, execution-local scheduler. Such placement implements the same publication and retirement transactions; it creates no program-visible queue, second completion condition, logical identity, or Mapping policy. Program-visible, multi-process, or device-side shared queues still require explicit atomic and memory-order actors plus a compatible Fabric consistency and coherence realization; DynamicWork does not infer one.

148

Dynamic-domain completion does not close a !dataflow.channel, emit EOS, or terminate an unrelated graph or thread. A caller may use the ordinary !dataflow.thread_token to order downstream work after quiescence. Concurrent consumers that must discover end-of-stream still require an explicit payload protocol or a future independently specified operation; no channel open/close/reset state is added.

155

For a breadth-first traversal of a rooted tree, the root node is the launch payload. Processing a node emits one dataflow.work.spawn per child in the tree's canonical child order, then yields. A leaf only yields. The collective token retires after the last descendant yields even if the implementation queue was temporarily empty between parent execution and child visibility. General graph BFS duplicate suppression and concurrent visited-set updates additionally require explicit atomic software operations and a compatible Fabric consistency and coherence realization; DynamicWork does not hide either requirement.

165

2.3 Static And Dynamic Thread Identity

167

The Canonical Dataflow entity catalog identifies a static root launch with a RootThreadLaunchRef; it does not assign an independent ID to the referenced dataflow.thread definition. The exact definition is recovered through the launch's Dataflow-owned callee relation.

172

A logical thread point is a derived value in that root-launch context:

174
LogicalThreadPoint =
    DensePoint(RootThreadLaunchRef, coordinate_tuple)
  | DynamicPoint(RootThreadLaunchRef, WorkItemId)
180

The exact launch-parameter environment supplies extents and admitted ordinary integer parameters used to interpret the point, but does not create another identity object. A dense coordinate tuple is body-visible through the trailing index arguments. It is never implicitly flattened; a program that needs a linear identifier computes one explicitly from coordinates and extents.

186

One static root launch may execute repeatedly. Runtime owns one transient ThreadDispatchOccurrenceId per concrete dispatch, and a concrete dense instance is (ThreadDispatchOccurrenceId, coordinate_tuple). The domain_instance component of a WorkItemId is the same dispatch occurrence, not a second counter or persistent identity. Occurrence IDs disappear after execution and are unavailable to Mapping. If program behavior or Mapping must distinguish repeated launches, the Structured Program must expose that distinction as a coordinate, launch parameter, or DynamicWork stable-item component rather than relying on a hidden epoch.

196

3. Source Induction Variables

198

Source lower bounds and steps are ordinary launch operands. The thread body reconstructs each source induction variable mechanically:

201
source_iv_d = lower_d + coord_d * step_d
coord_d in [0, extent_d)
206

This equation is program semantics. It supports dynamic lower bounds and steps without making source-loop bounds part of the thread ABI. The SCF optimizer must compute an extent that covers exactly the selected source iteration domain and must preserve overflow and signedness semantics required by the source program.

212

4. Derived Values And Memory Views

214

Values and memrefs cross a thread launch as ordinary typed operands and become matching ordinary definition arguments. Tiling, local ranges, subviews, address calculations, and explicit linearization use upstream MLIR operations such as affine, arith, and memref while the program remains in the SCF stage. Loom does not add a metadata-only passthrough operation.

220

The ownership optimizer decides whether a derived computation remains on the InstructionCore or enters a loom.spatial_region. A computation selected for the SpatialCore must mechanically lower to the canonical Dataflow actor surface; otherwise it stays outside the graph or makes that candidate non-finalizable. Analysis facts such as alias classes, access ranges, and memory footprints remain derived analyses unless they change program semantics.

228

5. Mapping Boundary

230

Logical coordinates, WorkItemId, designated work payload, and launch parameters are software facts. SystemMapping's B_thread relation consumes the legal logical domain to select an AccCore for each instance. The first dynamic profile exposes one typed stable execution class for its thread definition; all items share that Mapping selection while retaining distinct WorkItemId values. Payload, queue order, worker assignment, and ordinal path cannot select Mapping or invent another item identity. Event-relative ResourceUse separately owns occupancy and release. Neither relation may reinterpret logical identity as Cartesian hardware position.

240

For a dense domain, B_thread is evaluated over the exact RootThreadLaunchRef, coordinate tuple, and Dataflow-owned launch-parameter projection. For DynamicWork it uses the root launch plus the stable-item projection. The transient dispatch occurrence is deliberately absent: one verified relation applies to every execution of the same exposed logical domain.

247

Data partitioning visible to the program is expressed by its ordinary index and view computations. Physical placement of storage, memory services, and routes is owned by SpatialMapping and SystemMapping. No software view silently selects a fabric.pe, fabric.mem, transport endpoint, or protocol.

252

6. Verification And Tests

254

Anchor-level verification covers:

256
  • exact agreement between launch extent count and callee coordinate rank;
  • rank-zero, empty-domain, and negative-extent behavior;
  • exact ordinary-operand type agreement;
  • source-IV reconstruction for nonzero and dynamic lower or step values;
  • rejection of physical topology or Mapping authority in the software ABI;
  • root identity, deterministic child ordinals, acquire-before-publish, and exactly-once retirement for a dynamic work tree;
  • collective completion only when the dynamic responsibility set is empty; and
  • worker-independent item identity across Runtime assignment and stealing.
266

Tests should assert these stable boundaries rather than preserve a particular analysis cache, view-chain implementation, textual op order, or optimization heuristic.

270

7. Runtime Scheduling Boundary

272

The bounded scheduler profile is owned by docs/spec-runtime-abi.md. It may place not-yet-started responsibilities on finite-capacity worker deques, but it must preserve the same WorkItemId across assignments and delegate every root admission, spawn, and retirement transaction to DynamicWorkDomain. Queue pressure, worker idleness, cancellation requests, and scheduling trace state cannot complete the domain. The scheduler does not read SystemMapping; worker or deque selection cannot evaluate or replace B_thread.

280

The admitted execution adapter is narrower than the complete DynamicWork model. It accepts one root payload, no launch captures, and at most one direct graph launch. Dataflow projects the domain execution class; verified SystemMapping uses that class to select B_thread, B_graph, and contextual service plans for every item. The generic synchronous adapter may publish a finite child-payload group reported by its external execution owner and will drain the complete responsibility domain through bounded stealing. That report does not establish source-body execution or source publication lineage. The concrete CGRA entry further requires one byte-addressable scalar integer forwarded unchanged to the sole graph value input and a thread body containing only that launch and its yield. It derives graph runtime input from the scheduled payload, executes the selected SpatialMapping, and retires or cancels the item responsibility only after the selected execution returns. Captures, missing or non-direct graph bodies, nested or multiple graph launches, and incompatible payloads retain distinct typed capability reasons. Child publication has no canonical operation yet and therefore cannot enter a finalized program. Deployment image and hardware-provider transport for the stable-key relation remain a separately capability-gated contract.

299

8. Deferred Semantics

301

Program-visible or device-side shared work queues, priority queues, application duplicate suppression, distributed termination, active-item migration, and remapping remain deferred. They require explicit atomic, ordering, coherence, state-transfer, or service contracts. Distributed-buffer and neighborhood-exchange behavior likewise requires explicit dataflow and service semantics rather than hidden layout metadata.

308

Static halo and neighborhood exchange are compilation patterns, not new Dataflow entities. A structured candidate expresses them through existing logical-memory roots and views, explicit load/store dependencies, dense channel source maps, and ordinary thread/graph completion. Lowering may coalesce those operations only when it preserves the same visible memory, stream, and completion contract. No DistributedView, HaloArtifact, or implicit address ownership map is introduced.

316

The first version has no generic asynchronous bulk-movement operation. A provider-specific DMA or collective is admitted only after it has an independently observable software firing, completion, ordering, and failure contract that cannot be represented by the current operations. Until then, bulk copies lower to existing memory and channel semantics and use their exact completion tokens.

323

General device-side runtime spawn, spawn-then-feed of an independently blocked consumer, program-visible channel session/EOS operations, and arbitrary DynamicWork channel correspondence are also deferred. The bounded transient generation lifecycle in docs/spec-runtime-abi.md is permitted only for a complete finite dense channel invocation whose flat counts and endpoint membership are derived from existing launch correspondence. It does not add a Dataflow operation and is not inferred from dataflow.work.spawn or responsibility termination.