Nitin Keshav
Greater Sydney Area
2K followers
500+ connections
View mutual connections with Nitin
Nitin can introduce you to 10+ people at Technology by JLL
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Nitin
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
About
I’m a Lead Data Engineer with years of experience designing and scaling enterprise-grade…
Activity
2K followers
-
Nitin Keshav reposted thisNitin Keshav reposted thisGitHub listened! 🙌🙌🙌🙌🙌🙌 WE DID IT! GIthub Certifications now cost: $̶2̶0̶0̶ -> $99 USD Github Foundations $̶2̶0̶0̶ -> $99 USD Github Actions $̶2̶0̶0̶ -> $99 USD Github Administration $̶2̶0̶0̶ -> $99 USD Github Advanced Security Thanks to Github DA's for floating our message to whoever internally. [UPDATE] GitHub has a limited time coupon for 50% off the GitHub Foundations. Book it now, schedule your exam 3 weeks out. I will have a free study course for this to ace it. https://lnkd.in/grydjepZ #cloud #devops #git #aws #azure
-
Nitin Keshav shared thisTaking Orchestration game to next level. Prefect is perfect for bringing all workloads under one unified view. Although one can build the whole platform in Airflow over Kubernetes. With prefect customers pay a premium for one step deployment, managed workloads at hindsight and continous upgrades to match modern data stack. All in all happy to join the prefect certified community. next, hoping to ace the proof of value!Prefect Associate Certification • Nitin Keshav • Prefect • cHJvZHVjdGlvbjg2NzkyPrefect Associate Certification • Nitin Keshav • Prefect • cHJvZHVjdGlvbjg2Nzky
-
Nitin Keshav shared thisNitin Keshav shared this🚨Free Databricks Training & Certification Voucher Alert! 🏅if you want to get certified, here’s your chance ⬇️ We have 2 exciting GLOBAL trainings in October open to ALL #Databricks Prospects/Customers/Partners: ▪️Databricks #Lakehouse Fundamentals and the #Certification Overview Series. Register in Databricks Academy Catalog DB005a! (Links below) 📍Lakehouse Fundamentals Oct 18: APJ 12:00-14:30 (SGT) / EMEA 13:00-15:30 (BST) / AMER 09:30-12:00 (PT) This training focuses on the foundational concepts of Databricks Lakehouse Platform and will prepare you for the accreditation ▪️Customer/Prospect Registration: https://lnkd.in/eMJdbg7T ▪️Partner Registration: https://lnkd.in/ec_q-A5j 🏅Certification Overview Series - Associate In this training series, Databricks experts will share how to prepare for Associate-level Databricks certification exams 📍Data Engineer Associate Oct 19: APJ 12:00-15:00 (SGT) / EMEA 13:00-16:00 (BST) / AMER 9:30-12:30 (PT) ▪️Customer/Prospect Registration: https://lnkd.in/e73qYiMU ▪️Partner Registration: https://lnkd.in/eb88vjRT 📍Data Analyst Associate Oct 20: APJ 12:00-13:00 (SGT) / EMEA 13:00-14:00 (BST) / AMER 9:30-10:30 (PT) ▪️Customer/Prospect Registration: https://lnkd.in/eY2PFpP8 ▪️Partner Registration: https://lnkd.in/e2cPtBPW 📍Machine Learning Associate Oct 20: APJ 13:30-14:30 (SGT) / EMEA 14:30-15:30 (BST) / AMER 11:00-12:00 (PT) ▪️Customer/Prospect Registration: https://lnkd.in/ecKgQMTv ▪️Partner Registration: https://lnkd.in/euD5P2iA We will have a ton of TAs on to answer your questions. See you there! 🙏 #dataengineers #dataanalysts #mlpractitioners #freetraining #getcertified #levelup
-
Nitin Keshav shared thisNitin Keshav shared this𝗛𝗼𝘄 𝘁𝗼 𝘀𝘂𝗰𝗰𝗲𝗲𝗱 𝘁𝗵𝗿𝗼𝘂𝗴𝗵 𝗳𝗮𝗶𝗹𝘂𝗿𝗲? This is how great inventions were born! Failed experiments should be celebrated as ultimately they can lead to the innovations that change the world Thank you Simone Giertz for this demonstration of humility, learning and fun! ------------------------------- If you like my posts, you will enjoy my new book: https://lnkd.in/g4uCcg4 Click "Follow" for more #innovation insights https://lnkd.in/gFhhNg9 and https://lnkd.in/fjddMYP
-
Nitin Keshav shared thisNitin Keshav shared this- Forecasting: Principles and Practice - This is my #1 resource for time series forecasting. If you're interested in learning how to deal with time series data for data science, definitely check this one out! 👉 Free ebook -> http://dsdj.co/forecasting 👉 Free data science career training -> https://lnkd.in/gKE53aW #datascience
-
Nitin Keshav shared thisLove these guys!!Nitin Keshav shared this2 WAYS to improve your pie charts (everyone's favorite data-viz punching bag). First, ensure you're using them for the right reason. Second, practice! This month's #SWDchallenge is to find the perfect use case for a pie and wow us. #datastorytelling #dataviz #data READ "how to make a better pie chart": https://lnkd.in/e6uA7pe PRACTICE "perfect the pie": https://lnkd.in/e63iaC6
-
-
Nitin Keshav liked thisNitin Keshav liked thisWe’re growing our AI team at Zurich Australia and hiring across multiple levels as we continue our agentic AI journey. If you’ve got a strong, hands-on AWS + AI engineering background, check out these roles: • Lead AI Engineer – https://lnkd.in/gXPhaMS2 • Senior AI Engineer – https://lnkd.in/gNF5u7Zf • AI Domain Architect – https://lnkd.in/gy8YKDtp All roles are based in Sydney. If this sounds like you, or someone in your network, take a look and feel free to reach out. #Hiring #AgenticAI #AIEngineering #AWS #GenerativeAI
-
Nitin Keshav liked thisNitin Keshav liked thisMe, about a year ago. This was one of the last pictures I took at the office as a full-time employee. It was a period of the biggest transition in my career. I had a role that looked great on paper, but I knew I was ready for something different. I wanted to build something of my own, I just didn’t know what that would look like yet. I felt lost. Two months after this picture was taken, Casper, Eirini, and I started CAUCHY. By the end of the last year, I had left my full-time role. Our purpose is simple: we want to change the way AI consulting is done. We want to leave our customers in a better place than where we found them and with the knowledge to keep moving forward without us. Today, we have our own office in Amsterdam Zuidas, a growing team, and an exciting journey ahead. Looking back, I had no idea what was coming next. Sometimes, losing your sense of direction creates the space to find a path that truly feels like your own.
-
Nitin Keshav liked thisNitin Keshav liked this10 years later, I am back in the UK and Europe and this time, the journey feels very different. In 2016, I left Europe and moved to Australia. At the time, I was an continuing a career I had spent years building in insurance and technology. I certainly didn't imagine that, 10 years later, I would be back in this part of the world to establish and grow a business build from the ground up. But careers don't always follow the path you expect. After moving to Australia, Brian Siemsen and I started building what would eventually become Wilbur. We built the product, launched it in 2018, listened to our customers, learnt from the market, made plenty of changes along the way, and kept building. Over the years, what started in Australia has grown well beyond it. Today, Wilbur works with 55+ customers across Australia, New Zealand, the United States and South Africa. And now, exactly 10 years after moving away from Europe, Iam back in the UK and European market this time to write the next chapter for Wilbur. There is something quite special about that. I spent a significant part of my earlier career in this region, learning the insurance industry and working with insurers and technology teams. I definitely didn't expect that the journey would bring me back here with a product, a team, customers across multiple continents, and an opportunity to establish Wilbur in this market. The last decade has taught me a lot about insurance, technology, building products, customers, leadership, resilience and, probably most importantly, people. There have been successes, mistakes, difficult periods and plenty of learning. That's what makes looking back at the journey so rewarding. Australia became home. It also became the place where Wilbur was built. And now it feels like things have come full circle. Back in the UK and Europe after 10 years. Same industry. Familiar geography. A very different journey. And a new chapter for Wilbur. #Wilbur #Insurance #Claims #Insurtech #InsuranceTechnology #UKInsurance
-
Nitin Keshav liked thisPls reach out as we’d love to have you join our team as we continue to deliver the most critical and tangible outcomes with data and AI for our clients.Nitin Keshav liked this🚀 Sydney Data & AI Talent – Let's Connect We're growing our AI & Data team and are looking to connect with experienced professionals who are passionate about building modern data platforms and delivering AI-driven business outcomes across Databricks, Snowflake, AWS and Azure. I'm keen to speak with: 🔹 Databricks and Snowflake Platform Engineers 🔹 Data Engineers 🔹 Data Architects and Specialists Experience with PySpark, SQL, dbt, Unity Catalog, Terraform, cloud platforms, modern data architectures, AI and analytics is highly valued. Certifications are a strong plus. If you're passionate about solving complex data challenges, leading teams, working with clients, and delivering real business impact, I'd love to hear from you. 📩 Reach out for a confidential conversation. #Databricks #Snowflake #DataEngineering #DataArchitecture #AI #DataAndAI #AWS #Azure #dbt #PySpark #SydneyJobs #Hiring #Consulting #DataPlatforms Les Coleman-Stone Garima Purdhani Jay J. Ekta Nankani Raj Thakkar Aaron Wilkinson
-
Nitin Keshav liked this🚀 I’m Hiring! If your skills and experience align with this role, I’d love to hear from you! Feel free to reach out or share with someone who may be a great fit.Nitin Keshav liked thisIf you still think JLL is just a traditional real estate company, you're missing the massive tech engine we're building underneath. Through our Accelerate 2030 strategy, we're actively reshaping the entire commercial real estate industry. This isn't talk about digital transformation — we're putting serious resources into proprietary AI initiatives, predictive analytics, and smart building technologies. That's exactly why we're expanding our engineering presence in Jalisco, and why this Senior Cloud Engineer role is so critical to our Cloud team. Swaraj Gochhayat In this role, your work has a direct line of sight to our global business. You won't be managing servers or handling basic support tickets. You'll be building and scaling the modern cloud infrastructure that powers our AI and data platforms — the compute, networking, and automation layer that everything else runs on. Here's what you'll actually be doing: ✅ Design, build & support enterprise-scale GCP infrastructure (GKE, Compute Engine, Cloud Run, Cloud Functions, containers, databases), with working knowledge of Azure & AWS ✅ Monitor, troubleshoot & remediate GCP performance and security issues using Cloud Monitoring & Logging ✅ Partner with Dev teams to build and mature CI/CD pipelines primarily across GCP. ✅ Automate Infrastructure provisioning with Terraform, standardizing deployments across GCP ✅ If you're a senior cloud engineer in the Guadalajara area who wants to work on complex, high-impact systems where Cloud Engineering is treated as a critical discipline — not an afterthought — I'd love to connect. Apply: https://lnkd.in/g9fEgmvr
-
Nitin Keshav liked thisExcited to share that I have been granted a new U.S. patent🎉! in collaboration with my colleagues, Gaurav Jindal Neeraj Mantri Patent number: 12695699 Date of Patent: Jul 28, 2026 Patent Publication Number: 20250112863 This innovation designs a system that migrates gateways based on resource consumption. Explore the innovative solution: [Link to Patent Application] https://lnkd.in/de9kkTc3 #patent #innovation #colloboration #IntellectualProperty #Vmware
-
Nitin Keshav liked thisNitin Keshav liked thisI’m excited to share that I’ve joined the United Nations Operations and Crisis Centre (UNOCC) as a Data Engineering Intern! I’m looking forward to learning from experienced professionals, taking on new challenges, and contributing to meaningful, data-driven work throughout this journey. Grateful for this opportunity and excited for what lies ahead!
-
Nitin Keshav liked thisNitin Keshav liked thisAfter 2 years in the market, we are GA:ing our most requested capability: Unity AI Gateway: - Cost Controls on any model, harness, agent - Guardrails, Auditing, and ACLs on AI - Capacity for models from Anthropic, OpenAI, Gemini, Grok and Open Source Models Best part is that it’s all open source as part of MLflow! https://lnkd.in/g66R7MMRDatabricks: Build and Deploy Production-quality AI Agent SystemsDatabricks: Build and Deploy Production-quality AI Agent Systems
-
Nitin Keshav liked thisNitin Keshav liked this🚀 Thrilled to share that I've joined Lam Research as a Senior Data Analyst! 🎉 I'd like to thank meghana paalpare and Mohan Kumar for a smooth and well-organized onboarding experience. Looking forward to contributing to the team and growing in this new position. 💼✨ #DataAnalytics #SemiconductorIndustry #LamResearch #DataDriven
Experience
Education
-
PESIT (South)
-
-
Licenses & Certifications
-
Microsoft Certified: Azure Data Scientist Associate
Microsoft
IssuedCredential ID Certification Number: H345-1002 and Microsoft Certification ID: 989264646 -
Informatica PC Developer ICS exam 9.x
Informatica University
IssuedCredential ID Certificate No. 004-000134
Languages
-
Hindi
-
-
German
Limited working proficiency
-
English
Full professional proficiency
Recommendations received
5 people have recommended Nitin
Join now to viewView Nitin’s full profile
-
See who you know in common
-
Get introduced
-
Contact Nitin directly
Other similar profiles
Explore more posts
-
Abhisek Sahu
Rabobank • 177K followers
Lets Simplify Cloud Service Across Cloud AWS vs AZURE vs GCP ✨ Data Streaming - AWS + Kinesis = Real-Time Data Streaming - Azure + Event Hubs = Real-Time Data Streaming - GCP + Pub/Sub = Real-Time Data Streaming ✨ Data Warehousing - AWS + Redshift = Data Warehousing - Azure + Synapse Analytics = Data Warehousing - GCP + BigQuery = Data Warehousing ✨ Object Storage - AWS + S3 = Object Storage - Azure + Blob Storage = Object Storage - GCP + GCS = Object Storage ✨ Data Integration/Orchestration - AWS + Glue = Data Integration - Azure + Data Factory = Data Integration and Orchestration - GCP + Dataflow = Data Integration and Stream Processing ✨ Serverless Computing - AWS + Lambda = Serverless Computing - Azure + Functions = Serverless Computing - GCP + Cloud Functions = Serverless Computing ✨ Relational Databases - AWS + RDS = Managed Relational Databases - Azure + SQL Database = Managed Relational Databases - GCP + Cloud SQL = Managed Relational Databases ✨ NoSQL Databases - AWS + DynamoDB = NoSQL Databases - Azure + Cosmos DB = NoSQL Databases - GCP + Firestore = NoSQL Databases ✨ Big Data Processing - AWS + EMR (Elastic MapReduce) = Big Data Processing - Azure + HDInsight = Big Data Processing - GCP + Dataproc = Big Data Processing ✨ Machine Learning/AI - AWS + SageMaker = Machine Learning Platform - Azure + Machine Learning = Machine Learning Platform - GCP + AI Platform = Machine Learning Platform ✨ Content Delivery Network (CDN) - AWS + CloudFront = Content Delivery Network - Azure + Front Door/CDN = Content Delivery Network - GCP + Cloud CDN = Content Delivery Network ✨ Identity and Access Management - AWS + IAM = Identity and Access Management - Azure + Active Directory = Identity and Access Management - GCP + IAM = Identity and Access Management Image credits : Internet ! ♻️ Repost to help others grow 🔔 Follow Abhisek Sahu for more ♻️ I share cloud , data analysis/data engineering tips, real world project breakdowns, and interview insights through my free newsletter. 🤝 Subscribe for free here → https://lnkd.in/e8CaGD3p #engineering #cloud #aws #gcp #azure
416
46 Comments -
Raghavendra Betageri
Ninjacart • 853 followers
🚀 Deep Dive into HDFS: Building Distributed Storage from Scratch I've been on a fascinating journey exploring Hadoop HDFS fundamentals by building a local cluster (1 NameNode + 2 DataNodes) from the ground up. This hands-on experience has given me profound respect for the complexity of distributed systems. Why learn this in 2025? Today's cloud platforms (AWS EMR, Azure HDInsight, GCP Dataproc) provide managed solutions that abstract away infrastructure complexity. We can spin up Hadoop clusters with a few clicks—no manual configuration needed. But understanding what happens under the hood is invaluable. I have immense respect for Data Engineers and DevOps professionals who deployed these architectures in the pre-LLM era, armed only with documentation, StackOverflow threads, and YouTube tutorials. No ChatGPT to explain configs, no Copilot to autocomplete XML files—just pure grit, patience, and deep technical knowledge. What I've learned: ✅ How HDFS splits files into blocks and replicates them across DataNodes ✅ Fault tolerance mechanisms when nodes fail ✅ NameNode's role in metadata management vs DataNode's block storage ✅ Block size impact on storage efficiency and parallelism ✅ Administrative operations: quotas, snapshots, ACLs, replication management ✅ Integrating Python ETL workflows with HDFS as storage backend Key insight: HDFS alone provides distributed storage, NOT distributed compute. All processing in my setup happens client-side. This is just the foundation—YARN, MapReduce, Spark, and Hive build on top of this storage layer. Hands-on learning: I've documented everything in a GitHub repo—complete configs, step-by-step setup, experiments, and troubleshooting. Clone it, run it locally, and experiment! https://lnkd.in/gritdrRZ The repo uses small 10-byte blocks (vs 128MB production default) so you can visualise block splitting without needing gigabytes of test data. What's next: This is Part 1 of a series exploring the Hadoop ecosystem: HDFS (1NN + 2DN) ✅ Will explore YARN, Hive, Spark, Iceberg etc Want to go deeper? If you have access to multiple VMs, I highly encourage setting up each component on separate machines. Running NameNode and DataNodes on different physical/virtual hosts gives you authentic distributed system behavior—network latency, node failures, and all. Let's discuss! I'm eager to learn from the community. Drop your thoughts in the comments—happy to discuss, share, and learn together! #DataEngineering #Hadoop #HDFS #DistributedSystems #BigData #LearningInPublic #CloudComputing
14
1 Comment -
Hargovind Singh
Ericsson • 513 followers
𝗪𝗔𝗡𝗧 𝗔 ₹𝟯𝟱–𝟰𝟬+ 𝗟𝗣𝗔 𝗗𝗔𝗧𝗔 𝗘𝗡𝗚𝗜𝗡𝗘𝗘𝗥 𝗥𝗢𝗟𝗘? Then stop preparing like the interview is a Spark + SQL + Kafka trivia contest. Because at senior levels, the difficult part isn't knowing: repartition() ROW_NUMBER() consumer groups broadcast joins The difficult part is explaining WHEN, WHY, and WHAT HAPPENS WHEN IT FAILS. 🔵 𝗪𝗛𝗔𝗧 𝗦𝗘𝗡𝗜𝗢𝗥-𝗟𝗘𝗩𝗘𝗟 𝗜𝗡𝗧𝗘𝗥𝗩𝗜𝗘𝗪𝗦 𝗔𝗥𝗘 𝗥𝗘𝗔𝗟𝗟𝗬 𝗧𝗘𝗦𝗧𝗜𝗡𝗚 Imagine you're asked: → Your dashboard scans 3 TB for 24 hours. What do you change first? → A JOIN turns 100M rows → 4B rows. Do you tune Spark — or question the data model? → One Spark task is 12× slower than the median. Do you add more executors? → A broadcast join causes executor OOM. Was broadcasting actually the optimization? → Kafka processing succeeds, but the DB commit succeeds while offset commit fails. Can your pipeline safely recover? → Kafka lag keeps increasing. Why might adding consumers make zero difference? These aren't tool-knowledge questions. They're engineering judgment questions. 🟣 𝗧𝗛𝗘 𝗗𝗜𝗙𝗙𝗘𝗥𝗘𝗡𝗖𝗘 𝗜𝗦 𝗜𝗡 𝗧𝗛𝗘 𝗔𝗡𝗦𝗪𝗘𝗥 𝗠𝗶𝗱-𝗹𝗲𝘃𝗲𝗹: “Use Spark AQE.” 𝗦𝗲𝗻𝗶𝗼𝗿: “First I’ll inspect partition sizes, shuffle read/write, spill and key distribution. If the evidence shows skew, I’ll determine whether AQE can mitigate it or whether the data model/join strategy needs to change.” That difference is what separates knowing a technology from owning a system. 🟢 𝗜 𝗣𝗨𝗧 𝗧𝗛𝗜𝗦 𝗜𝗡𝗧𝗢 𝗔 𝗙𝗥𝗘𝗘 𝗙𝗜𝗘𝗟𝗗 𝗚𝗨𝗜𝗗𝗘 Inside: ▸ 16 architecture-grade interview problems ▸ SQL + Python + Spark + Kafka ▸ Production debugging framework ▸ System-design questions ▸ Failure & recovery patterns ▸ Cost vs latency trade-offs ▸ Replay + idempotency ▸ Engineering resources worth bookmarking ▸ 4-week architecture practice plan 𝗧𝗵𝗲 𝗴𝗼𝗮𝗹 𝗶𝘀𝗻’𝘁 𝘁𝗼 𝗺𝗲𝗺𝗼𝗿𝗶𝘇𝗲 𝟭𝟲 𝗮𝗻𝘀𝘄𝗲𝗿𝘀. It's to train yourself to think: Assumptions → Failure → Trade-off → Observability → Recovery 📘 𝗜’𝗩𝗘 𝗔𝗧𝗧𝗔𝗖𝗛𝗘𝗗 𝗧𝗛𝗘 𝗙𝗜𝗘𝗟𝗗 𝗚𝗨𝗜𝗗𝗘 I’ve attached a Senior Data Engineer Architecture & Interview Field Guide to help you prepare for senior-level interviews with real-world architecture and production scenarios. 📌 Save it. Practice it. Use it as your architecture drill. #DataEngineering #SeniorDataEngineer #SystemDesign #ApacheSpark #ApacheKafka
1
2 Comments -
Avinash S.
Mastercard • 20K followers
🚀 Snowflake Interview Questions for Senior Data Engineers – Part 2 In Part 1, we covered Architecture, Performance Optimization, and Data Modeling. Now, let’s dive into Security, Advanced Features, and Cost Optimization — areas where senior engineers are expected to think strategically. 🔹 4. Data Ingestion & Integration What are the different ways to load data into Snowflake (e.g., COPY, Snowpipe, external tables)? Can you explain how Snowpipe works and when it’s the right choice? How do external stages (S3, Azure Blob, GCS) integrate with Snowflake? What’s the role of file formats (Parquet, ORC, JSON, CSV) in Snowflake ingestion? How do you design a robust ELT/ETL pipeline into Snowflake? 🔹 5. Security & Governance How do Snowflake roles and grants work? What is the purpose of network policies in Snowflake? How do you enable row-level security (RLS) and column-level security (CLS)? Can you explain masking policies and their use cases? How do you implement data sharing securely across accounts or organizations? 🔹 6. Advanced Features How do streams and tasks enable near real-time pipelines in Snowflake? Can you explain the difference between replication and database failover in Snowflake? What is Time Travel in Snowflake, and how is it useful in recovery scenarios? How does zero-copy cloning work, and when would you use it? How would you use Snowflake’s external functions with AWS Lambda or Azure Functions? 🔹 7. Monitoring & Cost Management How do you monitor warehouse usage and query performance in Snowflake? What techniques do you use to optimize Snowflake cost without impacting performance? How do you analyze query history and warehouse utilization to detect inefficiencies? What are the best practices for managing multi-cluster warehouses? How do you track storage costs, including Time Travel and Fail-safe data? 👉 Thinking of making Part 3 all about real-world scenarios in Snowflake — like, “Your costs just doubled, how do you figure out why?” Should I go ahead and share those? Follow Avinash S. for more. #Snowflake #DataEngineering #CloudData #InterviewPrep #SeniorDataEngineer Would you like me to also create a Part 3 post that includes Scenario-Based Questions (e.g., “Your Snowflake costs doubled last month, how would you investigate?”) for a more practical angle?
65
3 Comments -
Sravani D
Morgan Stanley • 2K followers
How Databricks Stopped a Streaming App From “Freezing” During a New Episode Drop: A popular streaming platform dropped a new episode at 8 PM. Millions of users jumped in at the same time and suddenly: ❗ Some users couldn’t load the episode ❗ Others saw the wrong video quality ❗ A few got logged out randomly On the backend, nothing looked broken, but something felt off. The culprit? A slow user-session pipeline that updates: -device type -region -subscription status -streaming quality settings One micro-batch got delayed, the platform started serving old user-session data. Here’s how the team fixed it using Databricks: ✔ Used Databricks' job run view to detect the delayed batch. ✔ Used Delta Lake time travel to compare fresh vs stale session data. ✔ Enabled Auto Optimize and file compaction to speed up ingestion. ✔ Restarted the pipeline with proper schema enforcement. ✔ Added quality expectations to flag this next time. Within minutes: 🎥 Episodes loaded correctly ⚡ Streaming quality adjusted in real time 📈 User drop-offs went back to normal 😌 No major outage, just quick data engineering magic #DataEngineer #Databricks #Pyspark #DeltaLake #BigData #DataEngineering #RealTimeData #DataPipelines #ETL #DataQuality #Analytics #CloudComputing #Azure #AWS #GCP #SQL #AI #DataOps #Observability #C2C #USJobs #Hiring #ITTechJobs
4
1 Comment -
Kapil Rangwani
T-Mobile • 2K followers
I built a full Snowpipe setup runbook — here's everything you need to automate real-time data ingestion on Snowflake. Zero manual ingestion. 10 GB/sec. Sub-10s latency. The Snowpipe 2026 runbook is here. ❄️ Snowpipe is Snowflake's serverless auto-ingestion engine — it watches an S3 bucket, detects new files the moment they land, and loads them into Snowflake tables automatically, with zero manual intervention. I've packaged an end-to-end implementation runbook covering all 8 setup steps, annotated screenshots from a live environment, and the 10 most impactful Snowpipe Streaming enhancements released between 2025 and 2026. 🔹 4-Phase Architecture: The setup follows four sequential phases — Create File Format → Create Stage → Create Pipe (AUTO_INGEST=TRUE) → Configure S3 Event Notification with the Snowflake SQS ARN. Each phase must complete before the next, and the runbook walks through each with exact SQL and AWS console steps. 🔹S3 → SQS → Snowpipe Bridge: The critical link is the notification_channel ARN retrieved via DESC PIPE. This unique SQS ARN is pasted into the AWS S3 Event Notification, creating the live trigger. Recreating a pipe changes the ARN — a common production gotcha covered in the runbook. 🔹 10 GB/sec High-Performance Streaming (GA 2026): The new Snowpipe Streaming architecture delivers sub-10-second ingest-to-query latency at 10 GB/sec per table — a leap from the minutes-latency file-based classic model. Priced at a flat 0.0037 credits/GB, it eliminates complex per-core and per-file charges. 🔹 Auto Schema Evolution: Pipelines now automatically adapt when upstream data sources add new fields — no manual DDL, no downtime. This is a game-changer for fast-moving domains like 5G network data, IoT telemetry, and event streaming. 🔹Multi-Cloud GA (AWS + Azure + GCP): High-performance Snowpipe Streaming is now generally available across all cloud, removing vendor lock-in for hybrid and multi-cloud enterprises. 🔹Event Table Observability: Snowpipe now publishes detailed ingestion events — pipe state changes, per-file progress, and periodic digests — directly into Snowflake's Event Table, giving DevOps and SRE teams production-grade pipeline visibility. 🔹 Real-World Use Cases Covered: The runbook maps each concept to industry applications — Telecom streaming billions of network events, banks catching fraud before transactions complete, factories predicting equipment failures, ICU patient vitals feeding AI early-warning systems, and financial platforms querying pre-clustered time-series data at speed. Whether you're a data engineer building your first Snowflake pipeline or an architect designing petabyte-scale real-time infrastructure. Every step includes the exact command, the AWS console path, a screenshot from a live implementation, and a pro tip drawn from production experience. If this was helpful kindly repost and follow Kapil Rangwani for more insights!!
15
1 Comment -
Shishir Ranjan
WinWire • 1K followers
Spark DAG Stage Splitting — A Performance Tuning Perspective Spark splits a DAG into stages only at shuffle boundaries—this is not an implementation detail, it’s the core execution model. A new stage is created whenever data must move across partitions, which typically happens due to wide transformations such as: • Aggregations (groupBy, reduceByKey) • Non-broadcast joins • distinct operations • repartition and global sorting Why this matters: • Shuffles introduce network IO and disk spill • More stages increase scheduling overhead • Poor join strategy can explode stage count • Stage size directly impacts failure recovery time Key tuning levers highlighted in the visuals: • Prefer broadcast joins where possible • Use reduceByKey instead of groupByKey • Avoid unnecessary repartition • Push filters before joins • Cache only reused intermediates Understanding where and why stages are created is often the difference between a slow Spark job and an efficient one. 👉 Curious to hear from others: Which Spark transformation do you optimize first when performance degrades? #ApacheSpark #SparkPerformance #DataEngineering #BigData #PerformanceTuning
1
-
Julian Vecchio
Cognify Search • 28K followers
Is this holding you back from securing a Principal, Lead or Staff level Analytics Engineering role? I had an interesting conversation with a top Senior Analytics Engineer recently. . On paper, they look like a perfect candidate for Lead / Principal roles: ✅ 7+ years in data ✅ Strong SQL & data modelling fundamentals ✅ Modern data stack experience ✅ Solid software engineering practices ✅ Currently at a tier-one tech company You’d expect interviews to be straightforward. But they’re struggling to land roles at the level they’re aiming for. When we dug into it, the reason became clear. All of their best examples of impact came from previous companies. When we tried to use examples from their current role, it was much harder, because of the type of work they’re doing now. ▫️ They’re in an embedded analytics setup: ▫️ Working closely with stakeholders ▫️ Building & refactoring models, dashboards, reports. ▫️ Responding to business data requests All valuable work. But they’re not: ▫️ Designing core data models used across teams ▫️ Shaping the analytics platform ▫️ Driving standards, governance or platform improvements When companies hire Lead or Principal Analytics Engineers, they’re usually looking for people who have: Built or evolved the analytics platform Owned core data models Created reusable systems used across teams Taken analytics from 0 → 1 or 1 → 3 If your work is mostly business-as-usual analytics, you rarely get the chance to build those stories. So you end up with an impressive logo on your CV But if you are not doing impactful work – you may not be adding to your employability. If you're aiming for Lead / Principal roles, it’s worth asking yourself: Am I building systems or just delivering outputs? Am I improving the analytics platform or just operating within it? Would my recent work impress a principal-level hiring manager? Does this resonate with anyone?
27
8 Comments -
Ravena O
Do the AI • 96K followers
A RAG system without evaluation isn't really a system. It’s a guess with a vector database attached. Many teams build RAG as: LLM + Vector DB + Retrieval → Done. Then real users arrive. Messy PDFs. Vague queries. Missing context. Hallucinations. Production RAG needs more than that. Here’s the 9-layer RAG stack: L0 → Deployment Controls latency, scalability, and cost. AWS, Google Cloud, Groq. L1 → Data Extraction Bad parsing = bad retrieval. Tables, columns, and documents need to be extracted correctly. Docling, LlamaParse, Unstructured. L2 → Embeddings Converts content into searchable vectors. Hybrid search helps capture both semantic meaning and exact terms. Gemini Embedding, Voyage, Qwen3-Embedding. L3 → Vector DB Stores and retrieves relevant chunks. pgvector, Qdrant, Weaviate, Pinecone. L4 → Framework Orchestrates retrieval, workflows, and agents. LangGraph, LlamaIndex, Haystack. L5 → LLM Generates the final response. But a better LLM cannot fix poor retrieval. Claude, Llama, Qwen, Kimi. L6 → Memory Maintains context across conversations and sessions. Mem0, Zep, LangGraph. L7 → Evaluation Measures whether your RAG actually works. Faithfulness, context precision, answer relevancy. RAGAS, LangSmith, Opik. L8 → Alignment & Observability Guardrails, tracing, citations, and monitoring help catch failures before users do. The key takeaway: RAG isn't just LLM + Vector DB. It’s a system where every layer needs to be measured, monitored, and improved. Which layer has caused you the most pain — extraction, retrieval, generation, or evaluation? Save this for your next RAG or AI Agent project. CC: Rakesh Gohel
122
16 Comments
Explore collaborative articles
We’re unlocking community knowledge in a new way. Experts add insights directly into each article, started with the help of AI.
Explore More