Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
27 commits
Select commit Hold shift + click to select a range
4383cec
feat(reporting): add TimeWindowEvent for fixed-duration tumbling wind…
tombonfert Sep 30, 2026
4c20abf
refactor(query-engine): reuse SolverConfig canonical column names in …
tombonfert Sep 30, 2026
d75a1d0
refactor(query-engine): rename TumblingWindowsExpression to TimeWindo…
tombonfert Sep 30, 2026
e9ce811
docs(impulse): add TimeWindowEvent API reference and update event skills
tombonfert Sep 30, 2026
4dfcd41
fix(query-engine): normalize TimeWindowExpression.window_length to fl…
tombonfert Sep 30, 2026
c82a6c0
Merge branch 'main' into feature/time_window_event
tombonfert Sep 30, 2026
5b316f6
feat(reporting): compute TimeWindowEvent windows natively from contai…
tombonfert Oct 1, 2026
588df91
fix(reporting): reject non-finite window_length in TimeWindowEvent an…
tombonfert Oct 1, 2026
7ecc300
feat(query-engine, reporting): support TIMESTAMP container boundaries…
tombonfert Oct 1, 2026
ce0fdea
feat(reporting, query-engine): hash TimeWindowEvent windows by positi…
tombonfert Oct 7, 2026
9847744
feat(reporting, query-engine): support RAW channel data with TimeWind…
tombonfert Oct 7, 2026
7f02bba
feat(reporting, query-engine): validate TimeWindowEvent.max_windows_p…
tombonfert Oct 7, 2026
e9a1b66
refactor(tests): use itertools.pairwise for adjacent window assertions
tombonfert Oct 7, 2026
54a0ed6
refactor(reporting): move shared event metadata helpers to ContainerB…
tombonfert Oct 7, 2026
4f9aebe
test(reporting): harden TimeWindowEvent integration assertions and sc…
tombonfert Oct 7, 2026
4a92625
docs(query-engine, skills): clarify epoch_unit only converts containe…
tombonfert Oct 7, 2026
0a238e3
feat(query-engine, reporting): replace epoch_unit with channel_time_u…
tombonfert Oct 7, 2026
b73437e
feat(query-engine, reporting): add solver_config.container_time_unit …
tombonfert Oct 7, 2026
070281f
feat(query-engine, reporting): unify TimeWindowEvent window computati…
tombonfert Oct 7, 2026
7b5442b
feat(reporting, query-engine): hash TimeWindowEvent event_instance_id…
tombonfert Oct 7, 2026
b4ee4b7
refactor(query-engine, reporting): extract with_window_bounds from So…
tombonfert Oct 7, 2026
44129db
refactor(reporting): move as_dict, required_channels, and hash helper…
tombonfert Oct 7, 2026
b632d51
docs(query-engine, skills): clarify container_metrics.start_ts/stop_t…
tombonfert Oct 7, 2026
5a9df2c
fix(query-engine): cast window-bound conversion factor to long to pre…
tombonfert Oct 7, 2026
e6105d0
docs(reporting): document final-window rounding edge case for fractio…
tombonfert Oct 7, 2026
9032910
test(query-engine): use basic_narrow_db fixture for window_bounds tests
tombonfert Oct 7, 2026
e425ec0
fix(query-engine, reporting): return struct of starts/ends from windo…
tombonfert Oct 8, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 31 additions & 1 deletion docs/impulse/docs/config/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -170,6 +170,36 @@ Top-level fields on `SolverConfig`:
the `project_id` column (after column-name mapping) of every table it reads that carries one —
`container_tags` (if configured), `container_metrics`, and `channel_mapping` (if configured).
Omit it if you don't need project-level scoping; the solver does not require it.
- `channel_time_unit` (`"s"` | `"ms"` | `"us"` | `"ns"`, optional) and `channel_time_origin`
(`"epoch"` (default) | `"container_start"`): the time frame of the channel timestamps
(`tstart`/`tend`, or `timestamp` when `data_type = "RAW"`). `"epoch"` means absolute epoch
numbers; `"container_start"` means time relative to the container's start
(`container_metrics.start_ts`, e.g. seconds since the recording started). Channel timestamps are never converted; they must already be
numbers in that frame.

- `container_time_unit` (`"s"` | `"ms"` | `"us"` | `"ns"`, optional): the unit of **numeric**
`container_metrics.start_ts`/`stop_ts`, when it differs from the channels' unit. For example,
boundaries in epoch ms and channel samples in epoch µs need `container_time_unit = "ms"` and
`channel_time_unit = "us"`. Requires `channel_time_unit`; not allowed for `TIMESTAMP`
boundaries, which carry their own unit. Unset means the numeric boundaries are already in the
channels' unit.

These settings are only used by `TimeWindowEvent`, whose windows must lie in the channel time
frame. It derives its window bounds from `container_metrics.start_ts`/`stop_ts`:
- origin `"epoch"`: `TIMESTAMP` boundaries as epoch numbers in `channel_time_unit`; numeric
boundaries converted from `container_time_unit` to `channel_time_unit` (as they are when
`container_time_unit` is unset);
- origin `"container_start"`: `0` to `stop_ts - start_ts`, converted the same way.

`channel_time_unit` is required when the boundaries are `TIMESTAMP` columns; a report with a
`TimeWindowEvent` fails with a clear error until it is set. `TIMESTAMP_NTZ` and `DATE` boundaries
are not supported. Everything else sees the original `container_metrics.start_ts`/`stop_ts`:
`ContainerEvent`, `measurement_dimension`, container filters, and UDFs that request them via
`apply(..., container_metrics=[...])` (a `TIMESTAMP` arrives there as a `pd.Timestamp`).

The channel time frame is part of the definition hash of every `TimeWindowEvent` and of the
aggregations scoped to it, so changing it recomputes them over all containers in incremental
mode instead of mixing time frames in the gold tables.

Per-table sections (each a `TableConfig`):

Expand All @@ -192,7 +222,7 @@ Internal column names that mappings can target:
| `tstart`, `tend`| Sample interval start/end on the `channels` table (RLE) |
| `timestamp` | Raw sample timestamp on the `channels` table (RAW mode; encoded into `tstart`/`tend`) |
| `is_plausible` | Boolean plausibility flag on the `channels` table (RAW mode); consumed by `drop_implausible_data` |
| `start_ts`, `stop_ts` | Measurement start/stop epoch timestamps on the `container_metrics` table — referenced by `ContainerEvent` to derive event-fact start/end |
| `start_ts`, `stop_ts` | Measurement start/stop epoch timestamps on the `container_metrics` table — referenced by `ContainerEvent` and `TimeWindowEvent` to derive event-fact start/end. May be `TIMESTAMP` (see `channel_time_unit`) |
| `value` | Sample value (or attribute value on the EAV tag table) |
| `key` | Attribute key on the EAV `container_tags` table |
| `priority` | Tie-breaker column on the `channel_mapping` table |
Expand Down
8 changes: 8 additions & 0 deletions docs/impulse/docs/data_model/silver_layer_schema.md
Original file line number Diff line number Diff line change
Expand Up @@ -180,6 +180,14 @@ for human-readable display, `start_ts`/`stop_ts` for the gold
the epoch-typed pair). Populate whichever your queries and
`measurement_dimensions` config need.

`start_ts`/`stop_ts` may also be `TIMESTAMP` columns. To use them with a `TimeWindowEvent`, set
[`solver_config.channel_time_unit`](../config/configuration.md#solver-column-mappings-and-filters)
to the unit of the channel sample timestamps (`tstart`/`tend`, or `timestamp` in the raw format),
plus `channel_time_origin="container_start"` if those are relative to the container start, so
the window boundaries share the samples' time base. Numeric `start_ts`/`stop_ts` may use a different
epoch unit than the channel samples (e.g. ms boundaries, µs samples); then also set
`solver_config.container_time_unit`. The columns themselves are not converted.

:::

---
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -126,6 +126,21 @@ so that solver code can always reference the same constants.
override for the channel mapping (alias) table.
- `channels` (`TableConfig`): Column mappings and filters for the channel data table.
- `unit_conversion` (`TableConfig`): Column mappings and filters for the unit conversion table.
- `channel_time_unit` (`{"s", "ms", "us", "ns"} or None`): Time unit of the timestamps in the ``channels`` table (``tstart`` / ``tend``, or
``timestamp`` for RAW data). Only used to compute ``TimeWindowEvent`` windows in
that unit (see ``solvers.utils.window_bounds.with_window_bounds``); required when
``container_metrics`` ``start_ts`` / ``stop_ts`` are ``TIMESTAMP`` columns. Nothing
else is converted: channel timestamps, and the ``start_ts`` / ``stop_ts`` seen by
UDFs, ``ContainerEvent`` and ``measurement_dimension``, keep their original values.
- `channel_time_origin` (`{"epoch", "container_start"}`): Origin of the channel timestamps: absolute epoch (default), or relative to the
container's ``start_ts``. Like :attr:`channel_time_unit`, only used for the
``TimeWindowEvent`` windows.
- `container_time_unit` (`{"s", "ms", "us", "ns"} or None`): Unit of **numeric** ``container_metrics`` ``start_ts`` / ``stop_ts``, when it differs
from :attr:`channel_time_unit` (e.g. boundaries in epoch ms, channels in µs). Only
used to convert them into :attr:`channel_time_unit` for the ``TimeWindowEvent``
windows; requires :attr:`channel_time_unit`. Unset means the numeric boundaries are
already in the channels' unit. Not allowed for ``TIMESTAMP`` boundaries, which carry
their own unit.

#### from\_json

Expand Down Expand Up @@ -215,6 +230,30 @@ def start_ts_col() -> str
Internal column name for the measurement-start epoch timestamp on container_metrics.


#### window\_start\_col

```python
def window_start_col() -> str
```

Internal column name for the container start in the channel time frame.

Added by ``solvers.utils.window_bounds.with_window_bounds``; prefixed so it cannot
clash with a customer column.


#### window\_stop\_col

```python
def window_stop_col() -> str
```

Internal column name for the container stop in the channel time frame.

Added by ``solvers.utils.window_bounds.with_window_bounds``; prefixed so it cannot
clash with a customer column.


#### stop\_ts\_col

```python
Expand Down Expand Up @@ -457,3 +496,25 @@ def col_map() -> dict[str, str]
Short-key → internal-column-name mapping for UDFs and caches.


#### reject\_implausible\_channels\_filter\_in\_raw

```python
def reject_implausible_channels_filter_in_raw(is_raw: bool) -> None
```

Raise if an is_plausible channels filter is set in RAW mode.

Such a filter runs before raw encoding and bridges intervals across dropped
samples instead of splitting them; use drop_implausible_data instead. No-op
when not raw.


#### validate\_container\_time\_unit\_requires\_channel\_time\_unit

```python
def validate_container_time_unit_requires_channel_time_unit()
```

``container_time_unit`` converts into ``channel_time_unit``, so it needs one.


Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ ContainerEvent — an event spanning the full measurement container.
## ContainerEvent

```python
class ContainerEvent(Event)
class ContainerEvent(ContainerBoundaryEvent)
```

Event that treats the full measurement container as a single event instance.
Expand All @@ -33,18 +33,6 @@ Initialise a ContainerEvent.
- `desc` (`str`): Human-readable description.
- `attributes` (`dict`): Key-value metadata for the event.

#### get\_id

```python
def get_id() -> int
```

Return a unique identifier derived from the event name.

**Returns**:

`int`: Positive 32-bit integer identifier.

#### get\_expression

```python
Expand Down Expand Up @@ -86,30 +74,6 @@ so the name of the event is hashed.

`int`: Hash value representing the computation definition.

#### as\_dict

```python
def as_dict() -> dict
```

Return a dictionary representation of the event.

**Returns**:

`dict`:

#### as\_spark\_row

```python
def as_spark_row() -> Row
```

Return a Spark ``Row`` representation.

**Returns**:

`Row`:

#### determine\_events

```python
Expand Down Expand Up @@ -142,21 +106,3 @@ produces one event instance per container.

`DataFrame`: Spark DataFrame matching ``EVENT_INSTANCE_FACT_SCHEMA``.

#### determine\_metadata\_df

```python
def determine_metadata_df(cls, spark: SparkSession,
events: list[ContainerEvent]) -> DataFrame
```

Create a Spark DataFrame containing event metadata.

**Arguments**:

- `spark` (`SparkSession`): Active Spark session.
- `events` (`list of ContainerEvent`): List of ContainerEvent objects.

**Returns**:

`DataFrame`: Spark DataFrame matching ``EVENT_DIMENSION_SCHEMA``.

Original file line number Diff line number Diff line change
@@ -0,0 +1,160 @@
---
sidebar_label: time_window_event
title: impulse_reporting.events.time_window_event
---

TimeWindowEvent — splits each container into consecutive fixed-duration windows.


## TimeWindowEvent

```python
class TimeWindowEvent(ContainerBoundaryEvent)
```

Event that divides each measurement container into consecutive fixed windows.

Unlike ``ContainerEvent`` (one instance per container), a ``TimeWindowEvent`` emits one
event instance per fixed-duration slice, tiling the container's ``start_ts`` / ``stop_ts``
span with windows of length ``window_length``. The final slice is clamped to the
container end.

The event fact is computed from ``container_metrics`` alone (via


#### \_\_init\_\_

```python
def __init__(name: str,
window_length: float,
desc: str = None,
required_channels: list[str] = None,
attributes: Mapping[str, str] = None,
max_windows_per_container: int = MAX_WINDOWS_PER_CONTAINER)
```

Initialize a TimeWindowEvent object.

**Arguments**:

- `name` (`str`): Name of the event.
- `window_length` (`float`): Fixed window length, in the same time unit as the underlying timestamps
(e.g. milliseconds-since-epoch). Must be strictly positive and finite.
- `desc` (`str`): Description of the event.
- `required_channels` (`list of str`): List of required channels for the event. Informational; stored in the event
dimension table.
- `attributes` (`Mapping[str, str]`): Key-value metadata for the event. ``window_length`` is surfaced here
automatically (without overriding a user-supplied key).
- `max_windows_per_container` (`int`): Maximum number of windows per container (default 1,000,000). A container
exceeding it fails the report with an error naming the limit, which usually
means ``window_length`` is in the wrong unit for the boundaries. Not part of
the definition hash.

**Raises**:

- `ValueError`: If ``window_length`` is not strictly positive and finite, or
``max_windows_per_container`` is not a positive integer.

#### set\_channel\_time

```python
def set_channel_time(unit: str | None,
origin: str = "epoch",
container_unit: str | None = None) -> None
```

Record the channel time frame the windows are computed in.

Set by ``Report.add_event`` from the report's ``solver_config``. Stored on the
expression, whose string form feeds the definition hashes of this event and of the
aggregations scoped to it.

**Arguments**:

- `unit` (`str or None`): The report's ``solver_config.channel_time_unit``.
- `origin` (`str`): The report's ``solver_config.channel_time_origin`` (default ``"epoch"``).
- `container_unit` (`str or None`): The report's ``solver_config.container_time_unit``.

#### get\_expression

```python
def get_expression() -> TimeSeriesExpression | None
```

Get the time series expression associated with the event.

**Returns**:

`TimeSeriesExpression or None`: The time-window expression for the event.

#### get\_event\_type\_str

```python
def get_event_type_str() -> str
```

Get the event type string for TimeWindowEvent.

**Returns**:

`str`: Event type string.

#### determine\_definition\_hash

```python
def determine_definition_hash() -> int
```

Calculate definition hash for the time-window event.

Only includes the expression string, which encodes the attributes that affect the
event results: ``window_length`` and the channel time frame (``channel_time_unit``,
``channel_time_origin``, ``container_time_unit``; omitted while unset / default).
Resizing the window or changing the time frame therefore forces a full recompute in
incremental mode.

Excludes: name, description, required_channels, max_windows_per_container,
report_id

**Returns**:

`int`: Hash value representing the computation definition.

#### determine\_events

```python
def determine_events(
cls,
spark: SparkSession,
events: list[TimeWindowEvent],
*,
solved_df: DataFrame = None,
query: QueryBuilder = None,
solver: QuerySolver = None,
pre_filtered_containers_df: DataFrame = None) -> DataFrame
```

Extract the event fact table for the given list of TimeWindowEvent objects.

Resolves the matching containers via the solver's filter pipeline (like
``ContainerEvent``) and computes each event's windows natively from the
containers' ``start_ts`` / ``stop_ts`` in the channel time frame
(``solvers.utils.window_bounds.with_window_bounds``), so every filtered container
gets windows.
Each window becomes one event instance (``start_ts < end_ts``) whose
``event_instance_id`` hashes its boundaries. The solve uses the same window function
for scoped aggregations (see :func:`window_intervals_udf`), so the ids match.

**Arguments**:

- `spark` (`SparkSession`): Spark session for data processing.
- `events` (`list of TimeWindowEvent`): List of TimeWindowEvent objects to process.
- `solved_df` (`DataFrame`): Not used by TimeWindowEvent (kept for interface compatibility).
- `query` (`QueryBuilder`): Query builder with filters applied.
- `solver` (`QuerySolver`): Solver whose filter pipeline is used for container resolution.
- `pre_filtered_containers_df` (`DataFrame`): Pre-filtered containers for incremental processing.

**Returns**:

`DataFrame`: Spark DataFrame containing event instance facts.

3 changes: 2 additions & 1 deletion docs/impulse/docs/references/api/sidebar.json
Original file line number Diff line number Diff line change
Expand Up @@ -106,7 +106,8 @@
"items": [
"references/api/impulse_reporting/events/basic_event",
"references/api/impulse_reporting/events/container_event",
"references/api/impulse_reporting/events/sequence_of_events"
"references/api/impulse_reporting/events/sequence_of_events",
"references/api/impulse_reporting/events/time_window_event"
],
"label": "impulse_reporting.events",
"type": "category"
Expand Down
Loading
Loading