Subword pieces, BPE merges and their ties, parse trees and strict JSON, and a wheel that builds to the same hash twice.
Next: Classical Machine Learning.
#MT24Learning #NLP
Subword pieces, BPE merges and their ties, parse trees and strict JSON, and a wheel that builds to the same hash twice.
Next: Classical Machine Learning.
#MT24Learning #NLP
One build, two files: a wheel and an sdist. Built twice, same SHA-256. PyPI never reuses a filename. And as text "1.10" sorts before "1.9".
#MT24Learning #Python #OpenSource
One build, two files: a wheel and an sdist. Built twice, same SHA-256. PyPI never reuses a filename. And as text "1.10" sorts before "1.9".
#MT24Learning #Python #OpenSource
2 + 3 * 4: 14 by the grammar, 20 by the other tree. Trees grow 2, 5, 14, 42 per added phrase. And a duplicate JSON key passes silently unless you check.
#MT24Learning #Parsing #NLP
2 + 3 * 4: 14 by the grammar, 20 by the other tree. Trees grow 2, 5, 14, 42 per added phrase. And a duplicate JSON key passes silently unless you check.
#MT24Learning #Parsing #NLP
10 merges, 95 symbols down to 28. 6 of 10 steps were ties, and the tie rule changes how unseen words split. Capital letters and digits are edge cases too.
#MT24Learning #NLP #Tokenization
10 merges, 95 symbols down to 28. 6 of 10 steps were ties, and the tie rule changes how unseen words split. Capital letters and digits are edge cases too.
#MT24Learning #NLP #Tokenization
One phrase, three cuts: 23 characters, 2 words that both fall to [UNK], 7 subword pieces with IDs.
Same text can cost up to 15x more tokens in another language (Petrov et al. 2023).
#MT24Learning #NLP #Tokenization
One phrase, three cuts: 23 characters, 2 words that both fall to [UNK], 7 subword pieces with IDs.
Same text can cost up to 15x more tokens in another language (Petrov et al. 2023).
#MT24Learning #NLP #Tokenization
A pipeline where every row is accounted for, a split that can't leak (1.000 by row, 0.440 by group), and storage that refuses bad values and seals each file.
Next: Tokenization & Parsing.
#MT24Learning #DataQuality
A pipeline where every row is accounted for, a split that can't leak (1.000 by row, 0.440 by group), and storage that refuses bad values and seals each file.
Next: Tokenization & Parsing.
#MT24Learning #DataQuality
A plain SQLite table took 0.91, 'n/a' and 1.7 without complaint. STRICT plus a CHECK kept only 0.91. Then I sealed the saved file with SHA-256, and one edited digit changed the digest.
#MT24Learning #DataQuality #SQLite
A plain SQLite table took 0.91, 'n/a' and 1.7 without complaint. STRICT plus a CHECK kept only 0.91. Then I sealed the saved file with SHA-256, and one edited digit changed the digest.
#MT24Learning #DataQuality #SQLite
I split the same 1,000 rows two ways. By row, copies leaked into the test set and a nearest-neighbour model scored 1.000. By group it scored 0.440, close to a coin flip.
#MT24Learning #MachineLearning #DataScience
I split the same 1,000 rows two ways. By row, copies leaked into the test set and a nearest-neighbour model scored 1.000. By group it scored 0.440, close to a coin flip.
#MT24Learning #MachineLearning #DataScience
A spreadsheet had turned the gene MARCH1 into '1-Mar'. My pipeline caught it and set it aside with the 'n/a' row. Six rows in and three loaded. A second run left it at three. No key and a plain INSERT: six.
#MT24Learning #DataEngineering #DataQuality
A spreadsheet had turned the gene MARCH1 into '1-Mar'. My pipeline caught it and set it aside with the 'n/a' row. Six rows in and three loaded. A second run left it at three. No key and a plain INSERT: six.
#MT24Learning #DataEngineering #DataQuality
Sprint 16: a paradigm falls when a rival exists, not at the first anomaly.
Sprint 17: six classes of model failure, six different tests.
Sprint 18: a failure record keeps where the model broke and where it still holds.
#MT24Learning #PhilosophyOfScience
Sprint 16: a paradigm falls when a rival exists, not at the first anomaly.
Sprint 17: six classes of model failure, six different tests.
Sprint 18: a failure record keeps where the model broke and where it still holds.
#MT24Learning #PhilosophyOfScience
A failure record should say where the model still holds, not only where it broke. My pendulum record: within 1% up to 22°, 5% up to 50°, off by 18% at 90°. Feynman: "nature cannot be fooled."
#MT24Learning #ScientificMethod
A failure record should say where the model still holds, not only where it broke. My pendulum record: within 1% up to 22°, 5% up to 50°, off by 18% at 90°. Feynman: "nature cannot be fooled."
#MT24Learning #ScientificMethod
Six classes. One polynomial showed three: degree 1 too stiff; degree 9 at 0.073 on training and 0.41 on new points, then 482 past the data. Challenger flew at 31°F; the coldest earlier launch was 53°F.
#MT24Learning #MachineLearning
Six classes. One polynomial showed three: degree 1 too stiff; degree 9 at 0.073 on training and 0.41 on new points, then 482 past the data. Challenger flew at 31°F; the coldest earlier launch was 53°F.
#MT24Learning #MachineLearning
Continental drift, 1912 to 1968. The sea-floor magnetic stripes sat unexplained until spreading read them, then the field switched in about five years. Kuhn's point holds: an anomaly needs a rival.
#MT24Learning #PhilosophyOfScience #HistoryOfScience
Continental drift, 1912 to 1968. The sea-floor magnetic stripes sat unexplained until spreading read them, then the field switched in about five years. Kuhn's point holds: an anomaly needs a rival.
#MT24Learning #PhilosophyOfScience #HistoryOfScience
Cauchy means never settle. At 20 per group, significant hits averaged 0.83 for a true 0.5. Open surgery won both stone sizes and lost overall. Every number rerun in code.
#MT24Learning #Statistics #ExperimentalDesign
Cauchy means never settle. At 20 per group, significant hits averaged 0.83 for a true 0.5. Open surgery won both stone sizes and lost overall. Every number rerun in code.
#MT24Learning #Statistics #ExperimentalDesign
Simulated treatment, true effect +1. When doctors pick the sicker patients, the naive comparison says -1.262. Adjusting for severity: 0.996. A coin flip, no adjustment: 0.998.
#MT24Learning #ExperimentalDesign #Statistics
Simulated treatment, true effect +1. When doctors pick the sicker patients, the naive comparison says -1.262. Adjusting for severity: 0.996. A coin flip, no adjustment: 0.998.
#MT24Learning #ExperimentalDesign #Statistics
True effect d = 0.5, 10,000 simulated runs. 64 per group: power 0.794. 20 per group: power 0.335, and the significant runs averaged an effect of 0.831. Underpowered hits come out inflated.
#MT24Learning #Statistics #HypothesisTesting
True effect d = 0.5, 10,000 simulated runs. 64 per group: power 0.794. 20 per group: power 0.335, and the significant runs averaged an effect of 0.831. Underpowered hits come out inflated.
#MT24Learning #Statistics #HypothesisTesting
Running mean of 100,000 exponential draws: 0.949, 1.025, 1.0. Same test on a Cauchy: one value of 32,383 pushed the mean from 1.66 to 6.40, and it never settled. The law of large numbers needs a finite mean.
#MT24Learning #Probability #Statistics
Running mean of 100,000 exponential draws: 0.949, 1.025, 1.0. Same test on a Cauchy: one value of 32,383 pushed the mean from 1.66 to 6.40, and it never settled. The law of large numbers needs a finite mean.
#MT24Learning #Probability #Statistics
Gradients, logic, axioms and tests. The habit that linked them: before trusting a claim, run a check that could prove it wrong.
Next: Phase 5, Statistics & Experimental Design.
#MT24Learning #Mathematics #Logic
Gradients, logic, axioms and tests. The habit that linked them: before trusting a claim, run a check that could prove it wrong.
Next: Phase 5, Statistics & Experimental Design.
#MT24Learning #Mathematics #Logic
A property test states a rule for all inputs; the tool hunts for a case that breaks it. On floats, (a+b)+c == a+(b+c) broke fast. Dijkstra, 1969: testing shows "the presence of bugs, but never … their absence".
#MT24Learning #Python #SoftwareTesting
A property test states a rule for all inputs; the tool hunts for a case that breaks it. On floats, (a+b)+c == a+(b+c) broke fast. Dijkstra, 1969: testing shows "the presence of bugs, but never … their absence".
#MT24Learning #Python #SoftwareTesting
Euclid's parallel postulate resisted proof for two thousand years. Beltrami's 1868 model showed why: a geometry without it is as consistent as Euclid's own. Change that one axiom and triangles change.
#MT24Learning #Mathematics #Geometry
Euclid's parallel postulate resisted proof for two thousand years. Beltrami's 1868 model showed why: a geometry without it is as consistent as Euclid's own. Change that one axiom and triangles change.
#MT24Learning #Mathematics #Geometry
Wason's four-card task: E, K, 4, 7. The rule: vowel on one side, even number on the other. The most common pick is E and 4. The right pair is E and 7, because only they can break the rule.
#MT24Learning #Logic #CriticalThinking
Wason's four-card task: E, K, 4, 7. The rule: vowel on one side, even number on the other. The most common pick is E and 4. The right pair is E and 7, because only they can break the rule.
#MT24Learning #Logic #CriticalThinking