#USAddress
fastaddress

Fastaddress is a Python package that keeps the familiar usaddress API while moving its CRF runtime to Rust for much faster US address parsing. It delivers 11.3x higher single-core throughput, scales to 360K+ addresses per second on eight thread…

https://github.com/vinvomero/fastaddress
September 5, 2026 at 12:15 AM
i have used this model to predict which parts of a string might be an address, and I use that to help identify related companies, contracts and political donations in Illinois; I think it's pretty useful
GitHub - datamade/usaddress: :us: a python library for parsing unstructured United States address strings into address components
:us: a python library for parsing unstructured United States address strings into address components - datamade/usaddress
github.com
December 2, 2025 at 4:58 PM
So like, I use this, which is not a large model but is an example of using machine learning and probability

It classifies text, tagging parts which you can choose to use or not. And it's very helpful for messy addresses
GitHub - datamade/usaddress: :us: a python library for parsing unstructured United States address strings into address components
:us: a python library for parsing unstructured United States address strings into address components - datamade/usaddress
github.com
August 9, 2025 at 5:58 PM
agrc-usaddress 0.6.1 (Alpha)

Parse US (optimized for Utah by AGRC) addresses using conditional random fields

Unknown author
🏠Homepage
June 6, 2025 at 9:00 AM
Feeling sad about the world so of course I’m watching the #USAddress tonight 🥴 #SelfSabatoge
March 5, 2025 at 2:20 AM
Text Classification without context
I’m working on building a system to predict column headers for csv files based on the content of each column. The labels to find are First Name, Last Name, Company, Address 1, Address 2, City, State, and ZIP. I’ve successfully fine-tuned a variant of BERT on my labeled training data, but still get poor predictions from it. I’ve also tried just using the python package usaddress which performs a similar task, which was also unreliable. There’s two aspects of this task that I believe explain why these don’t perform well for me. First, the csv files I’ll be running predictions on won’t have a reliable column ordering, they can be mixed up. Second, these models rely heavily on context to make good predictions. So, if I pass an entire row from the CSV for a prediction, if the ordering is messed up I will get bad predictions due to the disordered context confusing the model. The other option I found was making predictions on each individual field, by asking it to predict a label for just strings like “1234 Main St” or “San Francisco” without including any other text from the row, but the models still don’t perform well on this, as they just have no context now, instead of the incorrectly ordered context. It seems that these types of models are overly complex for my task, and that the attention to context that normally helps these models make better predictions is actually what’s hurting my predictions here since I’m either depriving the model of context or giving it disordered context. I’m just looking for some insight into how to approach this kind of task, and/or if there are better choices of models to use for this type of task besides transformers since those all rely on context. I am still pretty new to machine learning, so apologies for any inaccuracies or if I didn’t explain this clearly enough.
discuss.huggingface.co
February 20, 2025 at 2:13 AM
I had to look it up because I forget, but it might have been this one. Its set up for shoppers, which I assume is where the money kicks in.

www.usaddress.com/get-usaddress
Get your FREE USaddress | USaddress.com
www.usaddress.com
January 31, 2024 at 12:32 AM