Mohammed Tahri
banner
mohammedtahri.bsky.social
Mohammed Tahri
@mohammedtahri.bsky.social
Systems Engineer · Neuro-Symbolic AI · Epistemic Evaluation Engines
Sprint 29/90: ridge regression on 65 terms, 5-fold CV over 15 penalties. α = 0.001: train 2,037, CV 4,027. Best α = 31.6: CV 3,047 ± 203. The one-standard-error rule picked α = 100; on 133 sealed patients it scored 3,717 vs 3,760. #MT24Learning #MachineLearning #Statistics
October 5, 2026 at 7:12 PM
Sprint 28/90: logistic regression, a decision tree and 5 nearest neighbours on 178 wines. In raw units k-NN scored 0.704: proline (278 to 1,680) swamped flavanoids (0.34 to 5.08). With z-scores, 0.944. The tree did not move. #MT24Learning #MachineLearning #DataScience
October 5, 2026 at 2:51 PM
Sprint 27/90: baselines first. Only 54 of 540 test digits are 9s, so "always no" scores 0.900 and finds none. Balanced accuracy: 0.500. A one-pixel rule picked by accuracy scored 0.894; picked by balanced accuracy, it found 47 of 54. #MT24Learning #MachineLearning #DataScience
October 4, 2026 at 6:53 PM
Sprint 26/90 opens Phase 9: supervised learning. I scored three rules on 30 held-out iris flowers. Memorising answered only 6 of them. Nearest neighbour and a single cut on petal width each missed 3, even though training error had ranked them 1 vs 3. #MT24Learning #MachineLearning #DataScience
October 4, 2026 at 1:25 PM
Phase 8 complete: Tokenization & Parsing, Sprints 22–25.

Subword pieces, BPE merges and their ties, parse trees and strict JSON, and a wheel that builds to the same hash twice.

Next: Classical Machine Learning.

#MT24Learning #NLP
October 3, 2026 at 3:48 PM
Sprint 25/90: Publishing a Package.

One build, two files: a wheel and an sdist. Built twice, same SHA-256. PyPI never reuses a filename. And as text "1.10" sorts before "1.9".

#MT24Learning #Python #OpenSource
October 3, 2026 at 2:03 PM
Sprint 24/90: Parsing Structured Meaning from Text.

2 + 3 * 4: 14 by the grammar, 20 by the other tree. Trees grow 2, 5, 14, 42 per added phrase. And a duplicate JSON key passes silently unless you check.

#MT24Learning #Parsing #NLP
October 3, 2026 at 1:40 AM
Sprint 23/90: Byte Pair Encoding & Its Edge Cases.

10 merges, 95 symbols down to 28. 6 of 10 steps were ties, and the tie rule changes how unseen words split. Capital letters and digits are edge cases too.

#MT24Learning #NLP #Tokenization
October 2, 2026 at 9:11 PM
Sprint 22/90: How Tokenization Works.

One phrase, three cuts: 23 characters, 2 words that both fall to [UNK], 7 subword pieces with IDs.

Same text can cost up to 15x more tokens in another language (Petrov et al. 2023).

#MT24Learning #NLP #Tokenization
October 2, 2026 at 5:46 PM
Phase 7 complete: Data Skepticism & Evaluation Data.

A pipeline where every row is accounted for, a split that can't leak (1.000 by row, 0.440 by group), and storage that refuses bad values and seals each file.

Next: Tokenization & Parsing.

#MT24Learning #DataQuality
October 1, 2026 at 11:31 PM
Sprint 21/90: validation and storage.

A plain SQLite table took 0.91, 'n/a' and 1.7 without complaint. STRICT plus a CHECK kept only 0.91. Then I sealed the saved file with SHA-256, and one edited digit changed the digest.

#MT24Learning #DataQuality #SQLite
October 1, 2026 at 7:57 PM
Sprint 20/90: building an evaluation dataset.

I split the same 1,000 rows two ways. By row, copies leaked into the test set and a nearest-neighbour model scored 1.000. By group it scored 0.440, close to a coin flip.

#MT24Learning #MachineLearning #DataScience
October 1, 2026 at 4:49 PM
Sprint 19/90: extract, transform, load.

A spreadsheet had turned the gene MARCH1 into '1-Mar'. My pipeline caught it and set it aside with the 'n/a' row. Six rows in and three loaded. A second run left it at three. No key and a plain INSERT: six.

#MT24Learning #DataEngineering #DataQuality
October 1, 2026 at 3:46 PM
Phase 6 summary: How Paradigms Die.
Sprint 16: a paradigm falls when a rival exists, not at the first anomaly.
Sprint 17: six classes of model failure, six different tests.
Sprint 18: a failure record keeps where the model broke and where it still holds.
#MT24Learning #PhilosophyOfScience
September 30, 2026 at 11:23 PM
Sprint 18/90: Documenting How a Model Failed.
A failure record should say where the model still holds, not only where it broke. My pendulum record: within 1% up to 22°, 5% up to 50°, off by 18% at 90°. Feynman: "nature cannot be fooled."
#MT24Learning #ScientificMethod
September 30, 2026 at 5:01 PM
Sprint 17/90: A Taxonomy of Model Failure.
Six classes. One polynomial showed three: degree 1 too stiff; degree 9 at 0.073 on training and 0.41 on new points, then 482 past the data. Challenger flew at 31°F; the coldest earlier launch was 53°F.
#MT24Learning #MachineLearning
September 30, 2026 at 4:28 PM
Sprint 16/90: Anatomy of a Paradigm Shift.
Continental drift, 1912 to 1968. The sea-floor magnetic stripes sat unexplained until spreading read them, then the field switched in about five years. Kuhn's point holds: an anomaly needs a rival.
#MT24Learning #PhilosophyOfScience #HistoryOfScience
September 30, 2026 at 2:59 PM
Phase 5 summary: Statistics & Experimental Design (Sprints 13–15).

Cauchy means never settle. At 20 per group, significant hits averaged 0.83 for a true 0.5. Open surgery won both stone sizes and lost overall. Every number rerun in code.

#MT24Learning #Statistics #ExperimentalDesign
September 29, 2026 at 11:11 PM
Sprint 15/90: Experimental Design & Rigorous Inference.

Simulated treatment, true effect +1. When doctors pick the sicker patients, the naive comparison says -1.262. Adjusting for severity: 0.996. A coin flip, no adjustment: 0.998.

#MT24Learning #ExperimentalDesign #Statistics
September 29, 2026 at 9:07 PM
Sprint 14/90: Hypothesis Testing, Power & Sample Size.

True effect d = 0.5, 10,000 simulated runs. 64 per group: power 0.794. 20 per group: power 0.335, and the significant runs averaged an effect of 0.831. Underpowered hits come out inflated.

#MT24Learning #Statistics #HypothesisTesting
September 28, 2026 at 5:51 PM
Sprint 13/90: Probability & Distributions.

Running mean of 100,000 exponential draws: 0.949, 1.025, 1.0. Same test on a Cauchy: one value of 32,383 pushed the mean from 1.66 to 6.40, and it never settled. The law of large numbers needs a finite mean.

#MT24Learning #Probability #Statistics
September 27, 2026 at 5:23 PM
Phase 4 of 27 is complete: Mathematical & Logical Foundations, Sprints 9–12.

Gradients, logic, axioms and tests. The habit that linked them: before trusting a claim, run a check that could prove it wrong.

Next: Phase 5, Statistics & Experimental Design.

#MT24Learning #Mathematics #Logic
September 27, 2026 at 12:54 AM
Sprint 12/90: testing and phase review.

A property test states a rule for all inputs; the tool hunts for a case that breaks it. On floats, (a+b)+c == a+(b+c) broke fast. Dijkstra, 1969: testing shows "the presence of bugs, but never … their absence".

#MT24Learning #Python #SoftwareTesting
September 26, 2026 at 12:39 AM
Sprint 11/90: axioms and declared assumptions.

Euclid's parallel postulate resisted proof for two thousand years. Beltrami's 1868 model showed why: a geometry without it is as consistent as Euclid's own. Change that one axiom and triangles change.

#MT24Learning #Mathematics #Geometry
September 25, 2026 at 9:16 PM
Sprint 10/90: logic, propositions and inference.

Wason's four-card task: E, K, 4, 7. The rule: vowel on one side, even number on the other. The most common pick is E and 4. The right pair is E and 7, because only they can break the rule.

#MT24Learning #Logic #CriticalThinking
September 25, 2026 at 5:09 PM