Skip to content

Latest commit

 

History

History
847 lines (712 loc) · 21 KB

File metadata and controls

847 lines (712 loc) · 21 KB

🔍 Passthrough Features

Passthrough Features in KDP

Handle IDs, metadata, and pre-processed data without unwanted transformations.

📋 Overview

Passthrough features allow you to include data in your model inputs without any preprocessing modifications. They're perfect for IDs, metadata, pre-processed data, and scenarios where you need to preserve exact values. With KDP v1.11.1+, you can now choose whether passthrough features are included in the main model output or kept separately for manual use.

🔑

ID Preservation

Keep product IDs, user IDs for result mapping

🏷️

Metadata Handling

Include metadata without processing it

⚡

Direct Integration

Include pre-processed data without modifications

🔄

Flexible Output

Choose between processed inclusion or separate access

🚀 When to Use Passthrough Features

🔑

IDs & Identifiers

Product IDs, user IDs, transaction IDs that you need for mapping results but shouldn't influence the model

🏷️

Metadata

Timestamps, source information, or other metadata needed for post-processing but not for ML

🔢

Pre-computed Features

Pre-computed embeddings or features that should be included in model processing

📊

Raw Values

Exact original values that need to be preserved without any transformations

🔍

Feature Testing

Compare raw vs processed feature performance in experiments

💡 Two Modes of Operation

KDP v1.11.1+ introduces two distinct modes for passthrough features:

🔄 Legacy Mode (Processed Output)

When: include_passthrough_in_output=True (default for backwards compatibility) Use case: Pre-computed features that should be part of model processing

preprocessor = PreprocessingModel(
    path_data="data.csv",
    features_specs=features,
    include_passthrough_in_output=True  # Default - backwards compatible
)

🎯 Recommended Mode (Separate Access)

When: include_passthrough_in_output=False Use case: IDs, metadata that should be preserved but not processed

preprocessor = PreprocessingModel(
    path_data="data.csv",
    features_specs=features,
    include_passthrough_in_output=False  # Recommended for IDs/metadata
)

📐 What the model returns in each mode

The two modes do not just move a column — they change the shape of the model's output, which is what your downstream code has to handle.

include_passthrough_in_output Model output
True (default) A single tensor. Passthrough columns are concatenated alongside the processed features, so they add to its width.
False A dict: {"processed": <tensor>, "passthrough": {<name>: <tensor>, ...}}. The processed tensor no longer contains the passthrough columns.
import tensorflow as tf

from kdp import FeatureType, PassthroughFeature, PreprocessingModel

features = {
    "amount": FeatureType.FLOAT_NORMALIZED,
    "row_id": PassthroughFeature(name="row_id", feature_type=FeatureType.PASSTHROUGH),
}
batch = {"amount": tf.constant([[50.0]]), "row_id": tf.constant([[7.0]])}

# Default: one tensor, two columns wide
combined = PreprocessingModel(path_data="data.csv", features_specs=features)
combined.build_preprocessor()
print(combined.model(batch).shape)          # (1, 2)

# Separate: the id is handed back untouched, next to the processed block
separate = PreprocessingModel(
    path_data="data.csv",
    features_specs=features,
    include_passthrough_in_output=False,
)
separate.build_preprocessor()
result = separate.model(batch)
print(result["processed"].shape)            # (1, 1)
print(result["passthrough"]["row_id"])      # the original value

💡 Defining Passthrough Features

from kdp import PreprocessingModel, FeatureType
from kdp.features import PassthroughFeature
import tensorflow as tf

# Simple approach using enum
features = {
    "product_id": FeatureType.PASSTHROUGH,  # Will use default tf.float32
    "price": FeatureType.FLOAT_NORMALIZED,
    "category": FeatureType.STRING_CATEGORICAL
}

# Advanced configuration with explicit dtype
features = {
    "product_id": PassthroughFeature(
        name="product_id",
        dtype=tf.string  # Specify string for IDs
    ),
    "user_id": PassthroughFeature(
        name="user_id",
        dtype=tf.int64   # Specify int for numeric IDs
    ),
    "embedding_vector": PassthroughFeature(
        name="embedding_vector",
        dtype=tf.float32  # For pre-computed features
    ),
    "price": FeatureType.FLOAT_NORMALIZED,
    "category": FeatureType.STRING_CATEGORICAL
}

🏗️ Real-World Example: E-commerce Recommendation

import pandas as pd
from kdp import PreprocessingModel, FeatureType
from kdp.features import PassthroughFeature
import tensorflow as tf

# Sample e-commerce data
data = pd.DataFrame({
    'product_id': ['P001', 'P002', 'P003'],      # String ID - for mapping
    'user_id': [1001, 1002, 1003],              # Numeric ID - for mapping
    'price': [29.99, 49.99, 19.99],             # ML feature
    'category': ['electronics', 'books', 'clothing'],  # ML feature
    'rating': [4.5, 3.8, 4.2]                   # ML feature
})

# Define features with proper separation
features = {
    # IDs for mapping - should NOT influence the model
    'product_id': PassthroughFeature(name='product_id', dtype=tf.string),
    'user_id': PassthroughFeature(name='user_id', dtype=tf.int64),

    # Actual ML features - should be processed
    'price': FeatureType.FLOAT_NORMALIZED,
    'category': FeatureType.STRING_CATEGORICAL,
    'rating': FeatureType.FLOAT_NORMALIZED
}

# Create preprocessor with separate passthrough access
preprocessor = PreprocessingModel(
    path_data="ecommerce_data.csv",
    features_specs=features,
    output_mode='dict',
    include_passthrough_in_output=False  # Keep IDs separate
)

model = preprocessor.build_preprocessor()

# Now you can:
# 1. Use the model for ML predictions (price, category, rating)
# 2. Access product_id and user_id separately for mapping results
# 3. No dtype concatenation issues between string IDs and numeric features

📊 How Passthrough Features Work

Passthrough Feature Model

Passthrough features create input signatures but can be processed or kept separate based on your configuration.

Legacy Mode Flow (include_passthrough_in_output=True)

➕

Input Signature

Added to model inputs with proper dtype

🔄

Minimal Processing

Type casting and optional reshaping only

🔗

Concatenated

Included in main model output (grouped by dtype)

Recommended Mode Flow (include_passthrough_in_output=False)

➕

Input Signature

Added to model inputs with proper dtype

🚫

No Processing

Completely bypasses KDP transformations

📦

Separate Access

Available separately for manual use

🔧 Configuration Options

Parameter Type Description
name str The name of the feature
feature_type FeatureType Set to FeatureType.PASSTHROUGH by default
dtype tf.DType The data type of the feature (default: tf.float32)
include_passthrough_in_output bool Whether to include in main output (True) or keep separate (False)

🎯 Advanced Examples

Example 1: Mixed Dtype IDs

import tensorflow as tf
from kdp import FeatureType, PassthroughFeature, PreprocessingModel

# Handles both string and numeric IDs without concatenation errors
features = {
    'product_id': PassthroughFeature(name='product_id', dtype=tf.string),
    'user_id': PassthroughFeature(name='user_id', dtype=tf.int64),
    'session_id': PassthroughFeature(name='session_id', dtype=tf.string),
    'price': FeatureType.FLOAT_NORMALIZED
}

preprocessor = PreprocessingModel(
    path_data="data.csv",
    features_specs=features,
    include_passthrough_in_output=False  # No dtype mixing issues
)

Example 2: Pre-computed Embeddings

import tensorflow as tf
from kdp import FeatureType, PassthroughFeature, PreprocessingModel

# Include pre-computed features in model processing
features = {
    'text_embedding': PassthroughFeature(
        name='text_embedding',
        dtype=tf.float32
    ),
    'image_embedding': PassthroughFeature(
        name='image_embedding',
        dtype=tf.float32
    ),
    'user_age': FeatureType.FLOAT_NORMALIZED
}

preprocessor = PreprocessingModel(
    path_data="embeddings.csv",
    features_specs=features,
    include_passthrough_in_output=True  # Include in model processing
)

Example 3: Metadata Preservation

import tensorflow as tf
from kdp import FeatureType, PassthroughFeature, PreprocessingModel

# Keep metadata for post-processing without affecting the model
features = {
    'timestamp': PassthroughFeature(name='timestamp', dtype=tf.string),
    'source_system': PassthroughFeature(name='source_system', dtype=tf.string),
    'batch_id': PassthroughFeature(name='batch_id', dtype=tf.int64),
    'sales_amount': FeatureType.FLOAT_NORMALIZED,
    'product_category': FeatureType.STRING_CATEGORICAL
}

preprocessor = PreprocessingModel(
    path_data="sales_data.csv",
    features_specs=features,
    include_passthrough_in_output=False  # Metadata separate from ML
)

⚠️ Important Considerations

⚠️

Dtype Compatibility

When using include_passthrough_in_output=True, passthrough features are grouped by dtype to prevent concatenation errors. String and numeric passthrough features are handled separately.

🎯

Recommended Usage

Use include_passthrough_in_output=False for IDs and metadata that shouldn't influence your model. Use True only for pre-processed features that should be part of model processing.

🔄

Backwards Compatibility

The default is True to maintain backwards compatibility with existing code. New projects should explicitly choose the appropriate mode.

🔍 Troubleshooting

❌ "Cannot concatenate tensors of different dtypes"

Solution: Set include_passthrough_in_output=False for ID/metadata features, or ensure all passthrough features have compatible dtypes.

❌ "inputs not connected to outputs"

Solution: This can happen with passthrough-only models. Ensure you have at least one processed feature, or use dict mode for passthrough-only scenarios.

❌ String features showing as tf.float32

Solution: Explicitly specify dtype in PassthroughFeature: dtype=tf.string

📈 Performance Tips

⚡

Reduce Model Complexity

Use include_passthrough_in_output=False for IDs to keep your model focused on actual ML features

🎯

Clear Separation

Separate concerns: IDs for mapping, features for ML, metadata for analysis

🔄

Choose the Right Mode

Legacy mode for pre-computed features, recommended mode for identifiers and metadata

<style> /* Base styling */ body { font-family: -apple-system, BlinkMacSystemFont, "Segoe UI", Roboto, Helvetica, Arial, sans-serif; line-height: 1.6; color: #333; margin: 0; padding: 0; } /* Feature header */ .feature-header { background: linear-gradient(135deg, #ff9800 0%, #f57c00 100%); border-radius: 10px; padding: 30px; margin: 30px 0; box-shadow: 0 4px 6px rgba(0,0,0,0.1); color: white; } .feature-title h2 { margin-top: 0; font-size: 28px; } .feature-title p { font-size: 18px; margin-bottom: 0; opacity: 0.9; } /* Overview card */ .overview-card { background-color: #fff; border-radius: 10px; padding: 20px 25px; margin: 20px 0; box-shadow: 0 2px 5px rgba(0,0,0,0.05); border-left: 4px solid #ff9800; } .overview-card p { margin: 0; font-size: 16px; } /* Key benefits */ .key-benefits { display: grid; grid-template-columns: repeat(auto-fill, minmax(250px, 1fr)); gap: 20px; margin: 30px 0; } .benefit-card { background-color: #fff; border-radius: 10px; padding: 20px; box-shadow: 0 4px 8px rgba(0,0,0,0.05); transition: transform 0.3s ease, box-shadow 0.3s ease; display: flex; flex-direction: column; align-items: center; text-align: center; } .benefit-card:hover { transform: translateY(-5px); box-shadow: 0 8px 16px rgba(0,0,0,0.1); } .benefit-icon { font-size: 2.5em; margin-bottom: 15px; } .benefit-card h3 { margin: 0 0 10px 0; color: #ff9800; } .benefit-card p { margin: 0; } /* Use cases */ .use-cases-container { display: grid; grid-template-columns: repeat(auto-fill, minmax(250px, 1fr)); gap: 20px; margin: 30px 0; } .use-case-card { background-color: #fff; border-radius: 10px; padding: 20px; box-shadow: 0 4px 8px rgba(0,0,0,0.05); transition: transform 0.3s ease, box-shadow 0.3s ease; display: flex; flex-direction: column; align-items: center; text-align: center; } .use-case-card:hover { transform: translateY(-5px); box-shadow: 0 8px 16px rgba(0,0,0,0.1); } .use-case-icon { font-size: 2.5em; margin-bottom: 15px; } .use-case-card h3 { margin: 0 0 10px 0; color: #ff9800; } .use-case-card p { margin: 0; } /* Architecture diagram */ .architecture-diagram { background-color: white; border-radius: 10px; padding: 20px; margin: 30px 0; box-shadow: 0 4px 8px rgba(0,0,0,0.05); text-align: center; } .architecture-image { max-width: 100%; border-radius: 5px; } .diagram-caption { margin-top: 20px; text-align: center; font-style: italic; } /* Workflow */ .workflow-container { display: grid; grid-template-columns: repeat(auto-fill, minmax(250px, 1fr)); gap: 20px; margin: 30px 0; } .workflow-card { background-color: #fff; border-radius: 10px; padding: 20px; box-shadow: 0 4px 8px rgba(0,0,0,0.05); transition: transform 0.3s ease, box-shadow 0.3s ease; display: flex; flex-direction: column; align-items: center; text-align: center; } .workflow-card:hover { transform: translateY(-5px); box-shadow: 0 8px 16px rgba(0,0,0,0.1); } .workflow-icon { font-size: 2.5em; margin-bottom: 15px; } .workflow-card h3 { margin: 0 0 10px 0; color: #ff9800; } .workflow-card p { margin: 0; } /* Code containers */ .code-container { background-color: #f8f9fa; border-radius: 8px; overflow: hidden; box-shadow: 0 2px 5px rgba(0,0,0,0.1); margin: 20px 0; } .code-container pre { margin: 0; padding: 20px; } /* Tables */ .table-container { margin: 30px 0; border-radius: 10px; overflow: hidden; box-shadow: 0 4px 8px rgba(0,0,0,0.05); } .config-table { width: 100%; border-collapse: collapse; } .config-table th { background-color: #fff3e0; padding: 15px; text-align: left; font-weight: 600; border-bottom: 2px solid #ff9800; } .config-table td { padding: 12px 15px; border-bottom: 1px solid #eaecef; } .config-table tr:nth-child(even) { background-color: #f8f9fa; } .config-table tr:hover { background-color: #fff3e0; } /* Considerations */ .considerations-container { display: grid; grid-template-columns: repeat(auto-fill, minmax(250px, 1fr)); gap: 20px; margin: 30px 0; } .consideration-card { background-color: #fff; border-radius: 10px; padding: 20px; box-shadow: 0 4px 8px rgba(0,0,0,0.05); transition: transform 0.3s ease, box-shadow 0.3s ease; display: flex; flex-direction: column; align-items: center; text-align: center; } .consideration-card:hover { transform: translateY(-5px); box-shadow: 0 8px 16px rgba(0,0,0,0.1); } .consideration-icon { font-size: 2.5em; margin-bottom: 15px; } .consideration-card h3 { margin: 0 0 10px 0; color: #ff9800; } .consideration-card p { margin: 0; } /* Best practices */ .best-practices-container { display: grid; grid-template-columns: repeat(auto-fill, minmax(250px, 1fr)); gap: 20px; margin: 30px 0; } .best-practice-card { background-color: #fff; border-radius: 10px; padding: 20px; box-shadow: 0 4px 8px rgba(0,0,0,0.05); transition: transform 0.3s ease, box-shadow 0.3s ease; display: flex; flex-direction: column; align-items: center; text-align: center; } .best-practice-card:hover { transform: translateY(-5px); box-shadow: 0 8px 16px rgba(0,0,0,0.1); } .best-practice-icon { font-size: 2.5em; margin-bottom: 15px; } .best-practice-card h3 { margin: 0 0 10px 0; color: #ff9800; } .best-practice-card p { margin: 0; } /* Navigation buttons */ .navigation-buttons { display: flex; justify-content: space-between; margin-top: 40px; } .nav-button { display: flex; align-items: center; padding: 12px 20px; background-color: #fff; border-radius: 8px; text-decoration: none; color: #333; box-shadow: 0 2px 5px rgba(0,0,0,0.05); transition: transform 0.3s ease, box-shadow 0.3s ease; } .nav-button:hover { transform: translateY(-2px); box-shadow: 0 4px 8px rgba(0,0,0,0.1); } .nav-icon { font-size: 1.2em; margin: 0 10px; } /* Responsive adjustments */ @media (max-width: 768px) { .key-benefits, .use-cases-container, .workflow-container, .considerations-container, .best-practices-container { grid-template-columns: 1fr; } .navigation-buttons { flex-direction: column; gap: 20px; } .nav-button { justify-content: center; } } </style>