#toymodel
It was a fun exercise to make Toymodel simulate its own logo within its own workflow. Since Toymodel is bounded-length Turing-complete, it can, at least in theory, simulate the core engine used to run Toymodel itself.

#sci-ml #toymodel #turing #animation #visualization #simulation
September 19, 2026 at 11:19 PM
In other news, my hubby got me the Super7 charging Godzilla from G-1!!!
#godzilla #figure #toymodel #bjd
January 29, 2025 at 5:34 AM
July 19, 2025 at 1:17 AM
This is the first in a series of posts introducing Toymodel (composable ML + optimization). What started as a project to build a financial optimization and prediction engine morphed into a new kind of scientific computing platform.

The public beta is out now.

More to follow.

toymodel.app
September 15, 2026 at 11:41 PM
Happy to have Rainer Engelken visiting us today for a great talk on analyzing neural network learning using techniques from dynamical systems.
January 22, 2025 at 10:01 PM
“D-dad?”
“Yesssss...” 🦖🚃🦖🚃🦖🚃
#powerrangers #mmpr #transformers #beastwars #toyphotography #hasbro #superf="/hashtag/supersentai" class="hover:underline text-blue-600 dark:text-sky-400 no-card-link">#supersentai #kiramager #megatron #yes #super #minipla #toymodel #actionfigurephotography #toys #hasbrotoypic #collection #collectibles #dinosaur #tyrannosaurusrex #train
November 26, 2024 at 3:47 PM
Don’t talk to me or my son again: Kiramager edition.
💎💎💎
#kirameiger #kiramager #superf="/hashtag/supersentai" class="hover:underline text-blue-600 dark:text-sky-400 no-card-link">#supersentai #toku #tokusatsu #powerrangers #super #sentai #minipla #toyref="/hashtag/toyphotography" class="hover:underline text-blue-600 dark:text-sky-400 no-card-link">#toyphotography #actionfigurephotography #collection #toy #toys #gem #mmpr #actionfigures #model #toymodel #hobby #scalemodel
November 26, 2024 at 2:58 PM
🔶Drill baby, drill.🔶

#kiramager #kirameiger #superf="/hashtag/supersentai" class="hover:underline text-blue-600 dark:text-sky-400 no-card-link">#supersentai #tokusatsu #powerrangers #super #minipla #toyref="/hashtag/toymodel" class="hover:underline text-blue-600 dark:text-sky-400 no-card-link">#toymodel #toys #toy #toycollector #toycommunity #toycollection #actionfigures #actionfigurephotography #bandai #sentai #orange #mashinsentaikiramager #mashinsentaikirameiger
November 25, 2024 at 7:11 PM
Getting zero attention in attention_module of Gemma3
Now, I am calling my python class to create the computation graph. ir = ToyModel(model, (input_ids, attention_mask), dynamic_shapes=dynamic_shapes) io_data = ir.predict(input_ids, attention_mask) ir.evaluation() function inside the model class to compute the things for attention is: def calculate_self_attention(Q, K, V, masked_fill=None, scale=None, epsilon=1e-9): """ Args: Q, K, V: [B, H, T_q, D] or [B, H, T_k, D] masked_fill: Optional additive mask of shape [B, 1, T_q, T_k] or [B, H, T_q, T_k] scale: Optional scaling factor (default: sqrt(D)) epsilon: Small constant for numerical stability """ print("Q dtype torch:", Q.dtype) print("K dtype torch:", K.dtype) print("V dtype torch:", V.dtype) B, H, T_q, D = Q.shape T_k = K.shape[2] # number of key tokens scale = scale or np.sqrt(D) log(f"Q: {np.sum(Q):.4f}, K: {np.sum(K):.4f}, V: {np.sum(V):.4f}") # Step 1: Raw attention logits QK_output = np.matmul(Q, K.transpose(0, 1, 3, 2)) # [B, H, T_q, T_k] logits_unmasked = QK_output / scale print(f"QK_output--- shape: {QK_output.shape}, value: {np.sum(QK_output):.4f}") print(f"logits_unmasked--- shape: {logits_unmasked.shape}, value: {np.sum(logits_unmasked):.4f}") ########################################################################### def _nz_stats(name, arr, tol=1e-12): total = arr.size zeros = np.count_nonzero(np.abs(arr) < tol) nonzeros = total - zeros pct = (nonzeros / total) * 100 print(f"{name}: nonzeros={nonzeros} ({pct:.2f}%), zeros={zeros}") # Debug: non-zero stats _nz_stats("Q: ", Q) _nz_stats("K: ", K) _nz_stats("V: ", V) _nz_stats("QK_output: ", QK_output) ########################################################################### # Step 2: Softmax over unmasked logits (for debugging or interpretability) A = np.exp(logits_unmasked - np.max(logits_unmasked, axis=-1, keepdims=True)) A = A / (np.sum(A, axis=-1, keepdims=True) + epsilon) log(f"A (unmasked attention weights) --- shape: {A.shape}, value: {np.sum(A):.4f}") # Step 3: Apply additive attention mask (optional) masked_fill = None if masked_fill is not None: logits_masked = logits_unmasked + masked_fill # [B, H, T, T] + [B, 1, T, T] log(f"masked_fill--- minimum: {np.min(masked_fill)}, maximum: {np.max(masked_fill)}") else: logits_masked = logits_unmasked.copy() log(f"logits_masked --- shape: {logits_masked.shape}, value: {np.sum(logits_masked):.4f}") # Step 4: Softmax over masked logits A_masked = np.exp(logits_masked - np.max(logits_masked, axis=-1, keepdims=True)) A_masked = A_masked / (np.sum(A_masked, axis=-1, keepdims=True) + epsilon) log(f"A_masked (masked attention weights)--- shape: {A_masked.shape}, value: {np.sum(A_masked):.4f}") # Step 5: Compute attention output using masked weights attention_output = np.matmul(A_masked, V) # [B, H, T_q, D] log(f"attention_output (using A_masked)--- shape: {attention_output.shape}, value: {np.sum(attention_output):.4f}") Output logs I am getting for `Gemma3` model: ToyModel: node='scaled_dot_product_attention_25', layer='Attention', func='scaled_dot_product_attention', parents='['clone_101', 'clone_102', 'clone_103', 'slice_494']', children='['transpose_105']' [DEBUG] number of inputs: 4 [DEBUG] idx: 0, item: torch.Size([1, 4, 2, 256]) [DEBUG] idx: 1, item: torch.Size([1, 4, 2, 256]) [DEBUG] idx: 2, item: torch.Size([1, 4, 2, 256]) [DEBUG] idx: 3, item: torch.Size([1, 1, 2, 2]) [DEBUG] Q: (1, 4, 2, 256), K: (1, 4, 2, 256), V: (1, 4, 2, 256), masked_fill: (1, 1, 2, 2) Q dtype torch: float32 K dtype torch: float32 V dtype torch: float32 [DEBUG] Q: 0.0000, K: -0.0000, V: 0.0000 QK_output--- shape: (1, 4, 2, 2), value: 0.0000 logits_unmasked--- shape: (1, 4, 2, 2), value: 0.0000 Q: : nonzeros=1013 (49.46%), zeros=1035 K: : nonzeros=1024 (50.00%), zeros=1024 V: : nonzeros=1024 (50.00%), zeros=1024 QK_output: : nonzeros=0 (0.00%), zeros=16 [DEBUG] A (unmasked attention weights) --- shape: (1, 4, 2, 2), value: 8.0000 [DEBUG] logits_masked --- shape: (1, 4, 2, 2), value: 0.0000 [DEBUG] A_masked (masked attention weights)--- shape: (1, 4, 2, 2), value: 8.0000 [DEBUG] attention_output (using A_masked)--- shape: (1, 4, 2, 256), value: 0.0000
discuss.huggingface.co
August 17, 2025 at 7:58 AM
Getting zero attention in attention_module of Gemma3
Output logs I am getting for `LLaMA3.2-1B` model: ToyModel: node='scaled_dot_product_attention_15', layer='Attention', func='scaled_dot_product_attention', parents='['clone_63', '_unsafe_view_30', '_unsafe_view_31', 'slice_311']', children='['transpose_64']' [DEBUG] number of inputs: 4 [DEBUG] idx: 0, item: torch.Size([1, 32, 2, 64]) [DEBUG] idx: 1, item: torch.Size([1, 32, 2, 64]) [DEBUG] idx: 2, item: torch.Size([1, 32, 2, 64]) [DEBUG] idx: 3, item: torch.Size([1, 1, 2, 2]) [DEBUG] Q: (1, 32, 2, 64), K: (1, 32, 2, 64), V: (1, 32, 2, 64), masked_fill: (1, 1, 2, 2) Q dtype torch: float32 K dtype torch: float32 V dtype torch: float32 [DEBUG] Q: 0.5828, K: 0.7185, V: -2.3097 QK_output--- shape: (1, 32, 2, 2), value: 0.1283 logits_unmasked--- shape: (1, 32, 2, 2), value: 0.0160 Q: : nonzeros=4096 (100.00%), zeros=0 K: : nonzeros=4096 (100.00%), zeros=0 V: : nonzeros=4096 (100.00%), zeros=0 QK_output: : nonzeros=128 (100.00%), zeros=0 [DEBUG] A (unmasked attention weights) --- shape: (1, 32, 2, 2), value: 64.0000 [DEBUG] logits_masked --- shape: (1, 32, 2, 2), value: 0.0160 [DEBUG] A_masked (masked attention weights)--- shape: (1, 32, 2, 2), value: 64.0000 [DEBUG] attention_output (using A_masked)--- shape: (1, 32, 2, 64), value: -2.3097
discuss.huggingface.co
August 17, 2025 at 7:58 AM
toymodelを卒業しないとな
February 24, 2025 at 10:27 AM
How brain pulsations drive solute transport in thecranial subarachnoid space: insights from a toymodel https://www.biorxiv.org/content/10.64898/2026.02.24.707684v1
February 26, 2026 at 2:48 AM
How brain pulsations drive solute transport in thecranial subarachnoid space: insights from a toymodel https://www.biorxiv.org/content/10.64898/2026.02.24.707684v1
February 26, 2026 at 2:48 AM