כדי לראות תפקידים מתאימים עליך להוסיף כישורים בפרופיל האישי במערכת COB.
ההרשמה והשימוש חינם!
מעולה, רוצה להירשם

הנדסה|מהנדס תוכנה|תוכנה
פורסם לפני יותר מחודשיים
פורסמה ברשת
Intelligent machines powered by Artificial Intelligence computers that can learn, reason and interact with people are no longer science fiction. GPU Deep Learning has provided the foundation for machines to learn, perceive, reason and solve problems. Today, visual computing is a crucial tool in helping people get along with technology, and we have extended its technology into datacenters, mobile devices and cars. There has never been a more exciting time to join our team – if this role sounds like a fit for you, we'd love to hear from you
we are seeking a Senior Performance Analysis Engineer to join our Performance group. In this exciting role, you will profile and analyze AI workloads on large GPUs and CPUs scale clusters for distributed Deep Learning LLM training focused on collectives communication and networking. You will interact with many types of hardware and platforms, such as HCAs, Switches, CPUs, GPUs, and Systems. You will develop performance analysis tools and methodologies to dive deeply into the details and understand performance expectations, limitations, and bottlenecks.
What you'll be doing:
Exploring and researching AI workloads and DL models specifically tailored for large-scale deep learning LLM training on our supercomputers and distributed systems focusing on high-performance networking and Nvidia Collective Communications Library (NCCL).
Benchmarking, Profiling, and Analyzing the performance to find bottlenecks and identify areas of improvement and optimizations, with a strong emphasis on networking aspects.
Implementing performance analysis tools.
Collaborating with many teams from hardware to software to provide performance analysis insights.
Defining performance test planning, setting performance expectations for new technologies and solutions, and working to reach the performance targets limits.

Requirements:
B.Sc. in Computer Science or Software Engineering or equivalent experience
5+ years of experience with high-performance Networking (RDMA, MPI, NCCL, Congestion Control Algorithms)
Demonstrated Performance Analysis skills and methodologies.
Experience with NVIDIA GPUs, CUDA library, deep learning frameworks like TensorFlow or PyTorch, combined with expertise in networking collective communication libraries (such as NCCL) and protocols (such as RoCE and RDMA).
Fast and self-learning capabilities with strong analytical and problem-solving skills.
Programming Languages: Python, Bash and C languages
Experience with Linux OS distros.
Great teammate with good communication and interpersonal skills
Ways to stand out from the crowd:
In-depth knowledge and experience with AI workloads and benchmarking for distributed LLM training.
Knowledge in CUDA, and NCCL libraries.
Knowledge in Congestion Control algorithms.
In-depth System knowledge and understanding (Intel / AMD / ARM CPUs, NVIDIA GPUs, HCA, Memory, PCI).
Strong Performance Analysis skills and methodologies using modern tools.

This position is open to all candidates.
משרות חדשות במערכת שיכולות לעניין אותך
פורסם לפני יותר מחודשיים
We are seeking a highly motivated Senior Diffusion Model Researcher to lead cutting-edge research in generative models with a focus ...
חולון / בת ים
פורסם לפני יותר מחודשיים
לאתר החברה בחולון דרוש.ה מהנדס.ת תוכנה rt Embedded מנוסה לאינטגרציה במעבדה. התפקיד כולל אחריות על פיתוח, תחזוקה ואינטגרציה של מערכות ...
פורסם לפני יותר מחודשיים
As a Senior Software Engineer on our Protocols team, youll work to take our Multi-Protocols Interface to the next level. ...
אזור מרכז - גוש דןתל אביב
פורסם לפני יותר מחודשיים
We are looking for a Sr Principal Linux Security Researcher for our Tel Aviv R&D center, to work on cortex-xdr ...
פורסם לפני יותר מחודשיים
Our team is looking for a Deep Learning Engineer.We are one of the few companies to have trained multi-billion parameter ...
פורסם לפני יותר מחודשיים
A cyber-security startup offering a fresh game-changing solution for data and access security. Our products create self-contained, cryptographically isolated workspaces which ...
חיפהנתניה
פורסם לפני יותר מחודשיים
A representative organization that puts target to bring & introduce new cutting edge products and technologies that will help local ...
נתניה
פורסם לפני יותר מחודשיים
החברה (אתר בנתניה) עוסקת בפיתוח, ייצור ואספקת פתרונות בדיקה חדשניים עבור מנועים חשמליים למגוון רחב של תעשיות ברחבי העולם במגוון ...
הצגת משרות נוספות
עדכון הכישורים שלך
להלן הכישורים הקיימים בפרופיל שלך. מומלץ להוסיף כישורים אשר דרושים למשרה או כישורים שלהערכתך רלוונטים לתפקיד.