##DataEngineering
Iniciando uma série sobre fundamentos de Engenharia de Dados!

No primeiro post, explico os conceitos básicos, desafios e o papel do engenheiro de dados no ecossistema moderno.

dev.to/1cadumagalha...

#DataEngineering #EngenhariaDeDados #BigData #ETL #Tech

RT do bem se puder
Introdução à Engenharia de Dados
Antes de começar com o roadmap, acho importante discutir o que realmente é Engenharia de...
dev.to
February 26, 2025 at 5:14 PM
I've released an update to SkaldMaps which improves the UI and fixes some data issues, based on early feedback from users: Filters, overlays, details, and more.

Check it out! https://skaldmaps.com

#RealEstate #PropertyResearch #DataEngineering
May 31, 2026 at 10:00 PM
September 22, 2026 at 6:43 AM
Armé un roadmap gratuito con conceptos básicos, desafíos técnicos y recursos sobre ingeniería de datos en español 🧙 en Github dónde pueden ir a darle estrellita, forkear y colaborar 🌟

#DataEngineering github.com/natayadev/da...
November 18, 2024 at 7:30 PM
After years in #dataengineering, I've noticed we focus too much on tools & tech, missing the fundamental challenges. It's not about picking the right stack—it's about mastering the complete data lifecycle.

The reality: We can't control source systems and their upstream data yet.
Challenges in Data Engineering
This chapter explores the multifaceted challenges within data engineering, providing insights into the data lifecycle, from collection to actionable insights. It delves into the complexities of data s...
dedp.online
January 7, 2025 at 5:15 PM
How can you make sure no one cares about your #OpenData, even if you are required to publish it? I wrote a short list of tips to help out: https://buff.ly/4hzPaN1 . Let me know if I missed something 😉. #DataScience #DataEngineering
How to Make Sure No One Cares About Your Open Data
Sharing data openly is a noble endeavor. It can drive research, innovation, and transparency. It is also really hard and annoying to do, plus you lose control - who knows what people will get up to.…
buff.ly
November 15, 2024 at 12:37 PM
Twitter feels like a post apocalyptic game, glad to have found an active community here!

#datatwitter #dataengineering
October 28, 2024 at 4:04 PM
PostgreSQL 19 is blazing fast. Your tools should be too.
Upgrade your workflow—not just your database.
Explore the PostgreSQL 19 companion suite built for speed.

valentina-db.com/en/postgresq...

#dataengineering #postgreSQL19
July 10, 2026 at 6:04 AM
I've created an additional resource called toolkit.ssp.sh.

The goal is to consolidate various #dataengineering skills from fundamental (Linux commands, containerization, programming languages) to Kubernetes orchestration.

The Toolkit provides the building blocks of data engineering work in 2025.
The Data Engineering Toolkit
The Data Engineering Toolkit To thrive as a data engineer, you need various skills—from fundamental (Linux commands, containerization, programming languages) to Kubernetes orchestration. The data eng...
toolkit.ssp.sh
June 20, 2025 at 11:26 AM
The most exciting data tool in recent times for me is @duckdb.org

It's very neat little package that packs a punch and you can basically run wherever you would like it. I went through DuckDB paper by @hannes.muehleisen.org and @markraasveldt.bsky.social

#databs #dataengineering
January 6, 2025 at 5:37 PM
Airflow for Beginners 🚀

A short tutorial by Sunjana Ramana for building an ETL data pipeline using Apache Airflow:
www.youtube.com/watch?v=3xyo...

#data #DataEngineering #airflow
Airflow for Beginners: Build Amazon books ETL Job in 10 mins
YouTube video by Sunjana in Data
www.youtube.com
December 21, 2024 at 2:54 PM
In my most recent blog I'm answering all the most important rhetorical #dataEngineering questions:

🫵 Why do *you* need CDC?
🤔 What even *is* CDC?
😝 Pffff, *I* don't need CDC (or do I?)
(with a bonus: why is ✨log-based CDC✨ the *best* way to do CDC? 🌶️)

👇
dcbl.link/why-cdc3
October 16, 2024 at 2:55 PM
We’re thrilled to have Emily Riederer keynoting at #positconf 2026! A leader at Capital One and a champion for open science, Emily is the expert on bridging the gap between data science and engineering.
✨ Level up your workflow with us: pos.it/conf
#rstats #pydata #DataEngineering #MachineLearning
March 16, 2026 at 4:03 PM
All the talks from the DuckCon conference 🦆 are now available:
www.youtube.com/playlist?lis...

#data #duckdb #dataengineering #datascience
July 4, 2026 at 10:35 PM
Time for @motherduck.com paper 📜

I've been a bit curious about @duckdb.org and MotherDuck( serverless duckdb offering) for while now.

I have been reading about it and getting hands dirty by building some toy examples with duckdb for my understanding

#dataBS #dataengineering #duckdb
January 19, 2025 at 5:09 PM
🚀 𝗖𝗮𝗻 𝗦𝗮𝗹𝗲𝘀𝗳𝗼𝗿𝗰𝗲 𝗗𝗮𝘁𝗮 𝗖𝗹𝗼𝘂𝗱 𝗛𝗮𝗻𝗱𝗹𝗲 𝗕𝗶𝗴 𝗗𝗮𝘁𝗮 𝗶𝗻 𝗥𝗲𝗮𝗹 𝗧𝗶𝗺𝗲?
📖 𝗥𝗲𝗮𝗱 𝘁𝗵𝗲 𝗳𝘂𝗹𝗹 𝗯𝗹𝗼𝗴: www.linkedin.com/feed/update/...

🎓 𝗟𝗲𝗮𝗿𝗻 𝘄𝗶𝘁𝗵 #Visualpath
🌐 𝗩𝗶𝘀𝗶𝘁: www.visualpath.in/data-cloud-t...

#Salesforce #SalesforceDataCloud #BigData #RealTimeData #DataCloud #Customer360 #DataEngineering #AI
Salesforce Data Cloud Training Ameerpet | Hands-On Projects | vamsi U
🚀 𝗖𝗮𝗻 𝗦𝗮𝗹𝗲𝘀𝗳𝗼𝗿𝗰𝗲 𝗗𝗮𝘁𝗮 𝗖𝗹𝗼𝘂𝗱 𝗛𝗮𝗻𝗱𝗹𝗲 𝗕𝗶𝗴 𝗗𝗮𝘁𝗮 𝗶𝗻 𝗥𝗲𝗮𝗹 𝗧𝗶𝗺𝗲? Modern businesses generate data from CRM systems, websites, apps, and customer interactions. But managing this data at scale requires the rig...
www.linkedin.com
September 22, 2026 at 7:44 AM
Let's talk about #apacheIceberg. It's been making waves in the #dataEngineering space, so you've likely heard of it. But what is it?

Iceberg is a high-performance, open table format designed for managing large-scale data workloads in a #dataLake. Now, why does that matter? 🧵
December 2, 2024 at 3:28 PM
This is genius. If I’m understanding correctly, metadata is handled by a DuckDB compatible SQL database and the actual data is handled by an open file format of your choice.

You can perform familiar SQL queries and DDL, on highly scalable open format data files. Well done! #databs #dataengineering
duckdb.org DuckDB @duckdb.org · May 27
Today we're launching DuckLake, an integrated data lake and catalog format powered by SQL. DuckLake unlocks next-generation data warehousing where compute is local, consistency central, and storage scales till infinity. ⁠ducklake is an open standard and we implemented it in the "ducklake" extension.
May 27, 2025 at 2:25 PM
Are you looking for something to learn in the coming break? My LinkedIn Learning course - Data Pipeline Automation with GitHub Actions Using R and Python- is open for a limited time 👇🏼

www.linkedin.com/posts/rami-k...

#Data #DataEngineering #DataScience #Rstats #Python
Rami Krispin on LinkedIn: #data #dataengineering #datascience #rstats #python
Are you looking for something to learn in the coming break? My LinkedIn Learning course - Data Pipeline Automation with GitHub Actions Using R and Python- is…
www.linkedin.com
December 25, 2024 at 4:36 PM
A short tutorial for setting a data lineage process with Airflow and Marquez by George Yates 👇🏼

📽️ www.youtube.com/watch?v=7cW-...

#dataengineering #data #airflow
How to Collect and Visualize Lineage Data from your Data Pipelines with Apache Airflow!
YouTube video by The Data Guy
www.youtube.com
November 17, 2024 at 5:25 AM
I'm a man of simplicity.

I don't know any other data stack that gets you from 0 to 1 as quickly...

Except Excel.

New vid 🎥: youtu.be/bbclf8ibIwM
#dataengineering #databs
1 YAML file is ALL your need to start your data stack
YouTube video by mehdio DataTV
youtu.be
December 16, 2024 at 5:23 PM
Ever wondered how to model topic 'trends' like Bluesky and others do? 📈 📉

I was curious too so I built a headline analytics pipeline which leverages a logistic growth model to do this! Orchestrated with @prefect.io and #dbt ⏳ 🔧

github.com/johnnyb1694/...

#datascience #dataengineering #python
GitHub - johnnyb1694/headline-analytics-pipeline: A data pipeline to extract & analyse data from publication headlines (e.g. the New York Times)
A data pipeline to extract & analyse data from publication headlines (e.g. the New York Times) - johnnyb1694/headline-analytics-pipeline
github.com
January 4, 2025 at 1:18 PM
🚀 Using Generated Columns in PostgreSQL 🚀

A generated column is a column with values calculated from other columns.

The example below calculates an order's discounted price from the order's total price.

#sql #dataengineering #dataanalytics #datascience
November 26, 2024 at 10:05 PM