dorsal/arxiv
View SchemaDepth-Wise Representation Development Under Blockwise Self-Supervised Learning for Video Vision Transformers
| Authors | Jonas Römer, Timo Dickscheid |
|---|---|
| Categories | |
| ArXiv ID | 2601.09040vv1 |
| URL | https://arxiv.org/abs/2601.09040 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
End-to-end backpropagation couples all layers through a global error signal, enabling coordinated learning but requiring long-range credit assignment. Motivated by recent progress in blockwise self-supervised learning (BWSSL), we ask whether masked video transformers can be trained without end-to-end backpropagation. Applying BWSSL to masked video modeling remains relatively underexplored and must handle spatiotemporal context and long-range temporal structure. More broadly, analyses that compare BWSSL and end-to-end training in terms of learning dynamics and depth-wise representation development remain sparse. We apply blockwise learning to a masked autoencoding video vision transformer by partitioning the encoder into blocks, each of which is optimized with a local masked reconstruction loss. Across model sizes and partition granularities, training converges and yields representations close to matched end-to-end baselines under linear-probe and retrieval proxies. In order to compare intermediate representations, we analyze depth-wise decodability, inter-block similarity, and patch-level diagnostics. Blockwise training exposes higher-level structure earlier, while later blocks saturate and operate in a more geometry-preserving regime. It can also induce token-level shifts consistent with stronger early mixing that pooled metrics can miss. These findings point to late-block saturation and interface formation as contributors to the remaining gap.
{
"annotation_id": "a927a18f-3a52-4326-98e1-1e50f7f06354",
"date_created": "2026-02-17T05:53:20.475000Z",
"date_modified": "2026-02-17T05:53:20.475000Z",
"file_hash": "3956eaa976980f4f7cd12907d091f4bc06db70eb4eff4666d9686eac75446ec4",
"private": false,
"record": {
"abstract": "End-to-end backpropagation couples all layers through a global error signal, enabling coordinated learning but requiring long-range credit assignment. Motivated by recent progress in blockwise self-supervised learning (BWSSL), we ask whether masked video transformers can be trained without end-to-end backpropagation. Applying BWSSL to masked video modeling remains relatively underexplored and must handle spatiotemporal context and long-range temporal structure. More broadly, analyses that compare BWSSL and end-to-end training in terms of learning dynamics and depth-wise representation development remain sparse. We apply blockwise learning to a masked autoencoding video vision transformer by partitioning the encoder into blocks, each of which is optimized with a local masked reconstruction loss. Across model sizes and partition granularities, training converges and yields representations close to matched end-to-end baselines under linear-probe and retrieval proxies. In order to compare intermediate representations, we analyze depth-wise decodability, inter-block similarity, and patch-level diagnostics. Blockwise training exposes higher-level structure earlier, while later blocks saturate and operate in a more geometry-preserving regime. It can also induce token-level shifts consistent with stronger early mixing that pooled metrics can miss. These findings point to late-block saturation and interface formation as contributors to the remaining gap.",
"arxiv_id": "2601.09040",
"authors": [
"Jonas R\u00f6mer",
"Timo Dickscheid"
],
"categories": [
"cs.CV"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Depth-Wise Representation Development Under Blockwise Self-Supervised Learning for Video Vision Transformers",
"url": "https://arxiv.org/abs/2601.09040",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "2076301c-91da-4ec2-989e-2e0be9436821",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}