<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0">
  <channel>
    <title>Case By Case</title>
    <link>https://stat-cbc.tistory.com/</link>
    <description>https://github.com/stat-eklee &amp;amp; https://blog.naver.com/2000051148 협업관련문의 : cowork.ek@gmail.com </description>
    <language>ko</language>
    <pubDate>Fri, 31 Jul 2026 20:22:51 +0900</pubDate>
    <generator>TISTORY</generator>
    <ttl>100</ttl>
    <managingEditor>통경</managingEditor>
    <image>
      <title>Case By Case</title>
      <url>https://tistory1.daumcdn.net/tistory/4339088/attach/f4b6c657f7c5434a84ed7af4e4d6c1bf</url>
      <link>https://stat-cbc.tistory.com</link>
    </image>
    <item>
      <title>[회고] 모두의 연구소 커리어랩 현직개발자 세미나 연사 후기</title>
      <link>https://stat-cbc.tistory.com/43</link>
      <description>&lt;p data-ke-size=&quot;size16&quot;&gt;대학원 석사과정을 졸업한 후, 몇 개의 회사를 전전하다 2022.09 부터 현재 회사에 재직하기 시작했습니다. 마침 좋은 기회를 받아, 커리어랩에서 진행하는 세미나에 연사로 참여할 수 있었습니다. (&lt;a href=&quot;https://modulabs.co.kr/careerlab/&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot;&gt;https://modulabs.co.kr/careerlab/&lt;/a&gt;)&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-origin-width=&quot;891&quot; data-origin-height=&quot;1260&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/pEEUl/btsjawNMQc7/45MJIzwk487HaikdOBuMSK/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/pEEUl/btsjawNMQc7/45MJIzwk487HaikdOBuMSK/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/pEEUl/btsjawNMQc7/45MJIzwk487HaikdOBuMSK/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FpEEUl%2FbtsjawNMQc7%2F45MJIzwk487HaikdOBuMSK%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;469&quot; height=&quot;663&quot; data-origin-width=&quot;891&quot; data-origin-height=&quot;1260&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;데이터 사이언티스트 / AI Researcher 두 직군을 모두 준비했었으나, 다른 두 연사분들께서 AI Researcher 직군에 있으셨기에 데이터 사이언티스트 직군에 대해 설명을 드렸습니다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;제가 준비한 부분은 크게 4가지로, 데이터 직무 소개/커뮤니티 활용해 커리어 성장하기 / JD(Job Description) 에 맞춘 내 이력 어필방법 / 면접후기를 준비했습니다. 실제 준비하면서 취업준비를 한 지 얼마 안되었었기에 솔직하고 다양한 경험을 공유드리려 노력했습니다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;데이터 직무 소개&lt;/h3&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;- 데이터 사이언티스트, 분석가, 엔지니어... 수 많은 직무 중에서 어떤 직무에 집중하고 이력서를 작성해야 할 지에 대해 고민이 많으실텐데, 이러한 직무별로 하는 일들이 다른 부분에 대해 설명드리고 본인이 어떤 특성을 가지고 있다면 이 직무가 맞겠다... 생각하실 수 있는 방향을 제시해드렸습니다. 실제 회사에서 일(프로젝트)이 주어졌을 때 직무별로 어떤 일을 수행해야하는지 예를 들어가면서 설명을 드리니 이해에 도움이 되신 것 같았습니다.&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;&lt;span style=&quot;color: #333333; text-align: start;&quot;&gt;커뮤니티 활용해 커리어 성장하기&lt;span&gt;&amp;nbsp;&lt;/span&gt;&lt;/span&gt;&lt;/h3&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;color: #333333; text-align: start;&quot;&gt;&lt;span&gt;- 어느 회사를 갈 진 모르겠지만, 데이터&amp;nbsp;사이언티스트가&amp;nbsp;될&amp;nbsp;거야.&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;- 나는 A회사의 데이터 사이언티스트가 될 거야.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;와 같이 case 를 구분해서 지금 뭐부터 해야 할 지에 대해 설명을 드렸습니다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;스크린샷 2023-06-08 오후 8.57.14.png&quot; data-origin-width=&quot;2174&quot; data-origin-height=&quot;1312&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/xiQ0q/btsjcWdCu8C/p9XqTp5z8yW3joAmvzMWWK/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/xiQ0q/btsjcWdCu8C/p9XqTp5z8yW3joAmvzMWWK/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/xiQ0q/btsjcWdCu8C/p9XqTp5z8yW3joAmvzMWWK/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FxiQ0q%2FbtsjcWdCu8C%2Fp9XqTp5z8yW3joAmvzMWWK%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;2174&quot; height=&quot;1312&quot; data-filename=&quot;스크린샷 2023-06-08 오후 8.57.14.png&quot; data-origin-width=&quot;2174&quot; data-origin-height=&quot;1312&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;color: #333333; text-align: start;&quot;&gt;그리고 이러한 필요한 커리어들을 커뮤니티 활동으로 어떻게 채우면 좋을지, 면접에서 커뮤니티에서 한 활동들을 어떻게 언급했는지도 정리해서 공유드렸습니다.&amp;nbsp;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;추가로, 면접 전에 참고했던 자료들도 링크를 공유드렸었습니다. &lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;&lt;a href=&quot;https://github.com/boostcamp-ai-tech-4/ai-tech-interview,&quot;&gt;https://github.com/boostcamp-ai-tech-4/ai-tech-interview,&lt;/a&gt; &lt;a href=&quot;https://zzsza.github.io/data/2&quot;&gt;https://zzsza.github.io/data/2&lt;/a&gt;&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;0&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;1&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;8&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;/&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;0&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;2/17/datascience-interivew-&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;q&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;uestions/&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;&lt;span style=&quot;color: #333333; text-align: start;&quot;&gt;&lt;span&gt;JD(Job Description) 에 맞춘 내 이력 어필방법&lt;/span&gt;&lt;/span&gt;&lt;/h3&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;color: #333333; text-align: start;&quot;&gt;&lt;span&gt;실제 필요한 역량, 우대사항, 개발환경을 확인하는 방법과 채용 프로세스나 인재상에 대해 확인하는 방법. 과제/코딩테스트 대비법에 대해 설명을 드리고 case 별로 if/else flow를 제시해서 어떤 상태이신 분은 이거부터 해라. 하고 말씀드렸습니다. 사실 제가 취업준비하면서 제일 필요하지만 어려워했던 건 당장 뭐부터 해야하지... 계획을 세우는 거였으니까요...!&amp;nbsp;&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;color: #333333; text-align: start;&quot;&gt;&lt;span&gt;상당수는 경험해봤어야 하는 것들, 당연히 할 줄 알아야 하는 것들... 로 구분해서 뭐부터 해야할지, JD 를 어떻게 해석해야 할 지 제시했습니다 : )&amp;nbsp;&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;color: #333333; text-align: start;&quot;&gt;&lt;span&gt;특히, 석사나 박사같은 학위가 필요한지에 대해서가 진짜 궁금하실 거 같아서 개인적인 정리를 했습니다.&amp;nbsp;&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;스크린샷 2023-06-08 오후 8.53.29.png&quot; data-origin-width=&quot;2286&quot; data-origin-height=&quot;1072&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/dHqf7k/btsi5Bh6DDI/9h4dHitRY7X1SvcT8iRymK/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/dHqf7k/btsi5Bh6DDI/9h4dHitRY7X1SvcT8iRymK/img.png&quot; data-alt=&quot;정말 간단히 적었지만, 정리한 결과만 봐도... 쉽지는 않다는 걸 아실 수 있을 겁니다...&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/dHqf7k/btsi5Bh6DDI/9h4dHitRY7X1SvcT8iRymK/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FdHqf7k%2Fbtsi5Bh6DDI%2F9h4dHitRY7X1SvcT8iRymK%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;2286&quot; height=&quot;1072&quot; data-filename=&quot;스크린샷 2023-06-08 오후 8.53.29.png&quot; data-origin-width=&quot;2286&quot; data-origin-height=&quot;1072&quot;/&gt;&lt;/span&gt;&lt;figcaption&gt;정말 간단히 적었지만, 정리한 결과만 봐도... 쉽지는 않다는 걸 아실 수 있을 겁니다...&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;면접후기&amp;nbsp;&lt;/h3&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;실제 면접에 참여해서 받았던 질문이나, 인상깊었던 부분 등에 대해 공유했습니다. 커피챗이나 자격증을 꼭, 어떤 걸 따야할 지에 대해서도 의견을 드렸습니다.&amp;nbsp;&lt;/p&gt;
&lt;div&gt;
&lt;div style=&quot;background-color: #ffffff;&quot;&gt;
&lt;div&gt;
&lt;div&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;- 경&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;험&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;했&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;던 &lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;면&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;접 &lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;유&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;형 키워&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;드 : 과제, 코딩테스트, 다대&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;일 &lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;면&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;접&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;, 다대다 면&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;접&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;, &lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;블&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;라인드 면&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;접&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;, 인&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;적&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;성, 실무&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;진 &lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;면&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;접&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;, &lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;임&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;원 면&lt;/span&gt;&lt;span style=&quot;font-family: NanumGothic; color: #5e5e5e;&quot;&gt;접 &lt;/span&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;[ 소개 ]&amp;nbsp;&lt;/h3&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;font-family: -apple-system, BlinkMacSystemFont, 'Helvetica Neue', 'Apple SD Gothic Neo', Arial, sans-serif; letter-spacing: 0px;&quot;&gt;커리어랩이란, 모두의연구소 커뮤니티를 활용해 &lt;/span&gt;&lt;span style=&quot;font-family: -apple-system, BlinkMacSystemFont, 'Helvetica Neue', 'Apple SD Gothic Neo', Arial, sans-serif; letter-spacing: 0px;&quot;&gt;기업의 성장을 돕는 채용 연계 서비스입니다. 실제로 커뮤니티에서 좋은 자리들이 있으면 커리어랩 담당자분이 등장하셔서 연계해주시는 거 같더라구요!&amp;nbsp; (&lt;a href=&quot;https://modulabs.co.kr/careerlab/&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot;&gt;https://modulabs.co.kr/careerlab/&lt;/a&gt;)&amp;nbsp;&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;스크린샷 2023-06-08 오후 9.07.32.png&quot; data-origin-width=&quot;904&quot; data-origin-height=&quot;332&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/cejnOr/btsjauCF0Sg/77S9yQthbydlk0RMCEkEcK/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/cejnOr/btsjauCF0Sg/77S9yQthbydlk0RMCEkEcK/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/cejnOr/btsjauCF0Sg/77S9yQthbydlk0RMCEkEcK/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FcejnOr%2FbtsjauCF0Sg%2F77S9yQthbydlk0RMCEkEcK%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;904&quot; height=&quot;332&quot; data-filename=&quot;스크린샷 2023-06-08 오후 9.07.32.png&quot; data-origin-width=&quot;904&quot; data-origin-height=&quot;332&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;[ 후기 ]&amp;nbsp;&lt;/h3&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;ChatGPT, Bard 등 수 많은 대화형 인공지능 모델들이 등장하고 하루에 몇 십, 몇 백 편씩 쏟아지는 논문의 홍수와 인공지능 업계의 전환점에 서있으면서 치열하게 회사생활을 하다보니 정신이 없다, 바쁘다는 핑계로 회고를 미뤄왔습니다. 앞으로는 글로 많은 경험들을 남기는 것을 소중하게 생각하고 싶네요,... 솔직하고 꼼꼼히 준비했던 만큼, 후기도 좋게 받을 수 있었던 거 같습니다.&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;edited_스크린샷 2023-06-08 오후 9.11.31.png&quot; data-origin-width=&quot;673&quot; data-origin-height=&quot;572&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/7or6u/btsjasrlkHI/vKVK50SVXyTS9LkWdJwSck/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/7or6u/btsjasrlkHI/vKVK50SVXyTS9LkWdJwSck/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/7or6u/btsjasrlkHI/vKVK50SVXyTS9LkWdJwSck/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2F7or6u%2FbtsjasrlkHI%2FvKVK50SVXyTS9LkWdJwSck%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;673&quot; height=&quot;572&quot; data-filename=&quot;edited_스크린샷 2023-06-08 오후 9.11.31.png&quot; data-origin-width=&quot;673&quot; data-origin-height=&quot;572&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;edited_스크린샷 2023-06-08 오후 9.02.35.png&quot; data-origin-width=&quot;1358&quot; data-origin-height=&quot;586&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/bA5unX/btsi5A4BLBo/UYiXVstgl3iDjPQVXzRN10/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/bA5unX/btsi5A4BLBo/UYiXVstgl3iDjPQVXzRN10/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/bA5unX/btsi5A4BLBo/UYiXVstgl3iDjPQVXzRN10/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FbA5unX%2Fbtsi5A4BLBo%2FUYiXVstgl3iDjPQVXzRN10%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;1358&quot; height=&quot;586&quot; data-filename=&quot;edited_스크린샷 2023-06-08 오후 9.02.35.png&quot; data-origin-width=&quot;1358&quot; data-origin-height=&quot;586&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;21세기에 가장 섹시한 직업, 데이터 사이언티스트로써 강의할 수 있어서 영광이었습니다~&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;스크린샷 2023-06-08 오후 8.58.22.png&quot; data-origin-width=&quot;1458&quot; data-origin-height=&quot;1102&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/wlMGX/btsjbqsMgFf/v0h21PgKkU1ExnPqwxgblk/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/wlMGX/btsjbqsMgFf/v0h21PgKkU1ExnPqwxgblk/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/wlMGX/btsjbqsMgFf/v0h21PgKkU1ExnPqwxgblk/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FwlMGX%2FbtsjbqsMgFf%2Fv0h21PgKkU1ExnPqwxgblk%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;1458&quot; height=&quot;1102&quot; data-filename=&quot;스크린샷 2023-06-08 오후 8.58.22.png&quot; data-origin-width=&quot;1458&quot; data-origin-height=&quot;1102&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;</description>
      <category>회고</category>
      <category>모두의연구소</category>
      <category>취업특강</category>
      <category>커리어랩</category>
      <author>통경</author>
      <guid isPermaLink="true">https://stat-cbc.tistory.com/43</guid>
      <comments>https://stat-cbc.tistory.com/43#entry43comment</comments>
      <pubDate>Thu, 8 Jun 2023 21:04:41 +0900</pubDate>
    </item>
    <item>
      <title>[후기] 삼성케어플러스 Z플립3 올갈이 후기</title>
      <link>https://stat-cbc.tistory.com/42</link>
      <description>&lt;p data-ke-size=&quot;size16&quot;&gt;어느덧 Z플립3을 출시하자마자 사전예약해 산지 1년이 지났다. 그 동안 떨어뜨린 기억이 없는데 어디서 자꾸 흠집이 생겼고, 어짜피 사전예약 혜택이 삼성케어플러스 1년 무료가입이었기에 이게 만료되기 전에 올갈이(메인보드 제외, 디스플레이-배터리-전면/후면부를 모두 수리하는 것)를 한 후기를 남기려고 한다.&amp;nbsp;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-origin-width=&quot;960&quot; data-origin-height=&quot;485&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/cX5FkM/btrK1pVHesI/9WqgOmOEz7rUShFE7VL6v0/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/cX5FkM/btrK1pVHesI/9WqgOmOEz7rUShFE7VL6v0/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/cX5FkM/btrK1pVHesI/9WqgOmOEz7rUShFE7VL6v0/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FcX5FkM%2FbtrK1pVHesI%2F9WqgOmOEz7rUShFE7VL6v0%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;960&quot; height=&quot;485&quot; data-origin-width=&quot;960&quot; data-origin-height=&quot;485&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;방문 전!! 꼭 확인해야 할 사항이 있다. 가는 AS지점에 부품이 모두 있다면 상관이 없지만, Z플립 특성상 파손이 잦아 다른 분들이 먼저 받았을 경우 부품 재고가 없는 경우가 있다. &lt;span style=&quot;color: #ee2323;&quot;&gt;&lt;b&gt;그렇기 때문에 꼭 1588-3366 으로 전화해서 Z플립3 의 색상과 AS 방문 센터 지점을 얘기하고 미리 예약해야 한다.&amp;nbsp;&lt;/b&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;예약을 하게 되면, 상담원 분이 도착에 최대 일주일 이상 걸린다고 안내를 해주신다. 그리고 도착하면 카카오톡으로 알람이 오는데, 나는 19일 금요일에 전화해서 자재 예약을 했었다(방문해서 재고가 없다는 점을 알고,,, 2번 방문했다)&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;blob&quot; data-origin-width=&quot;589&quot; data-origin-height=&quot;1440&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/CV6EP/btrKZ09Wy3s/CNv8xY1LyVF31tFKLdSfP1/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/CV6EP/btrKZ09Wy3s/CNv8xY1LyVF31tFKLdSfP1/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/CV6EP/btrKZ09Wy3s/CNv8xY1LyVF31tFKLdSfP1/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FCV6EP%2FbtrKZ09Wy3s%2FCNv8xY1LyVF31tFKLdSfP1%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;589&quot; height=&quot;1440&quot; data-filename=&quot;blob&quot; data-origin-width=&quot;589&quot; data-origin-height=&quot;1440&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;그리고 24일 수요일에 도착했다는 메시지를 받았다! 하지만 금요일에 휴가였기 때문에 금요일에 방문했다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;blob&quot; data-origin-width=&quot;589&quot; data-origin-height=&quot;1440&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/GZGMy/btrK4CzYpdo/RxyUxrQGxMbHeeV4Sv2tE0/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/GZGMy/btrK4CzYpdo/RxyUxrQGxMbHeeV4Sv2tE0/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/GZGMy/btrK4CzYpdo/RxyUxrQGxMbHeeV4Sv2tE0/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FGZGMy%2FbtrK4CzYpdo%2FRxyUxrQGxMbHeeV4Sv2tE0%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;589&quot; height=&quot;1440&quot; data-filename=&quot;blob&quot; data-origin-width=&quot;589&quot; data-origin-height=&quot;1440&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;Z플립3 를 올갈이 하면 총 14만원을 개인 부담금으로 결제하면 된다. 수리 시간은 30분 내외로 소요됐다! 자재를 예약했으니 이름을 말씀드리면 내가 예약한 자재를 사용해주신다. 수리기사님이 굉장히 친절하게 응대해주셨었다. 동생과 함께 가서 얘기하다보니 금방 시간이 갔다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;+ 삼성디지털프라자에서 사전예약하신 분들은 이 개인부담금 14만원도 환급해주는 이벤트를 했다고 하니 해당되시면 확인해보시면 될 거 같다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;나는 이마트에서 사서 해당이 안됐다. 그래도 배터리 같은 거 생각한다면 14만원 내고 충분히 교체할만 하다! 원가는 40얼마라고 한다... (덜덜)&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-origin-width=&quot;1080&quot; data-origin-height=&quot;1440&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/EhE4N/btrK0HbeZyl/z6LKcyOdfQshBiQa7kF4XK/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/EhE4N/btrK0HbeZyl/z6LKcyOdfQshBiQa7kF4XK/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/EhE4N/btrK0HbeZyl/z6LKcyOdfQshBiQa7kF4XK/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FEhE4N%2FbtrK0HbeZyl%2Fz6LKcyOdfQshBiQa7kF4XK%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;1080&quot; height=&quot;1440&quot; data-origin-width=&quot;1080&quot; data-origin-height=&quot;1440&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;다시 깔끔해진 내 폰... 앞으로도 오래오래 기스없이 살아야 한다~~&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;</description>
      <category>Z플립수리후기</category>
      <category>갤럭시Z플립</category>
      <category>삼성케어플러스</category>
      <category>삼케플</category>
      <author>통경</author>
      <guid isPermaLink="true">https://stat-cbc.tistory.com/42</guid>
      <comments>https://stat-cbc.tistory.com/42#entry42comment</comments>
      <pubDate>Wed, 31 Aug 2022 17:24:19 +0900</pubDate>
    </item>
    <item>
      <title>[Python] H2O 패키지를 활용하여 XGBoost 모델 구축하기</title>
      <link>https://stat-cbc.tistory.com/41</link>
      <description>&lt;p data-ke-size=&quot;size16&quot;&gt;최근 프로젝트를 수행하면서 H2O 라는 좋은 패키지를 활용해볼 수 있는 기회가 있어서 이에 대한 내용을 정리해보려고 한다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;a href=&quot;https://docs.h2o.ai/h2o/latest-stable/h2o-docs/welcome.html&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot;&gt;https://docs.h2o.ai/h2o/latest-stable/h2o-docs/welcome.html&lt;/a&gt;&lt;/p&gt;
&lt;figure id=&quot;og_1661927369122&quot; contenteditable=&quot;false&quot; data-ke-type=&quot;opengraph&quot; data-ke-align=&quot;alignCenter&quot; data-og-type=&quot;website&quot; data-og-title=&quot;Welcome to H2O 3 &amp;mdash; H2O 3.36.1.4 documentation&quot; data-og-description=&quot;Docs &amp;raquo; Welcome to H2O 3 Edit on GitHub Welcome to H2O 3 H2O is an open source, in-memory, distributed, fast, and scalable machine learning and predictive analytics platform that allows you to build machine learning models on big data and provides easy pro&quot; data-og-host=&quot;docs.h2o.ai&quot; data-og-source-url=&quot;https://docs.h2o.ai/h2o/latest-stable/h2o-docs/welcome.html&quot; data-og-url=&quot;https://docs.h2o.ai/h2o/latest-stable/h2o-docs/welcome.html&quot; data-og-image=&quot;https://scrap.kakaocdn.net/dn/bq4M4Q/hyPEJ7cV48/akpDpBYouTpBIjXBv2Rq8k/img.png?width=2790&amp;amp;height=596&amp;amp;face=0_0_2790_596,https://scrap.kakaocdn.net/dn/b0Fa03/hyPEMv5NcV/K84spxE4T8osZEPuL8thKk/img.png?width=2175&amp;amp;height=398&amp;amp;face=0_0_2175_398,https://scrap.kakaocdn.net/dn/crf0DS/hyPEEdKPmK/2RNKwaT2AWBQ6zsVMQ96wK/img.png?width=720&amp;amp;height=540&amp;amp;face=0_0_720_540&quot;&gt;&lt;a href=&quot;https://docs.h2o.ai/h2o/latest-stable/h2o-docs/welcome.html&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; data-source-url=&quot;https://docs.h2o.ai/h2o/latest-stable/h2o-docs/welcome.html&quot;&gt;
&lt;div class=&quot;og-image&quot; style=&quot;background-image: url('https://scrap.kakaocdn.net/dn/bq4M4Q/hyPEJ7cV48/akpDpBYouTpBIjXBv2Rq8k/img.png?width=2790&amp;amp;height=596&amp;amp;face=0_0_2790_596,https://scrap.kakaocdn.net/dn/b0Fa03/hyPEMv5NcV/K84spxE4T8osZEPuL8thKk/img.png?width=2175&amp;amp;height=398&amp;amp;face=0_0_2175_398,https://scrap.kakaocdn.net/dn/crf0DS/hyPEEdKPmK/2RNKwaT2AWBQ6zsVMQ96wK/img.png?width=720&amp;amp;height=540&amp;amp;face=0_0_720_540');&quot;&gt;&amp;nbsp;&lt;/div&gt;
&lt;div class=&quot;og-text&quot;&gt;
&lt;p class=&quot;og-title&quot; data-ke-size=&quot;size16&quot;&gt;Welcome to H2O 3 &amp;mdash; H2O 3.36.1.4 documentation&lt;/p&gt;
&lt;p class=&quot;og-desc&quot; data-ke-size=&quot;size16&quot;&gt;Docs &amp;raquo; Welcome to H2O 3 Edit on GitHub Welcome to H2O 3 H2O is an open source, in-memory, distributed, fast, and scalable machine learning and predictive analytics platform that allows you to build machine learning models on big data and provides easy pro&lt;/p&gt;
&lt;p class=&quot;og-host&quot; data-ke-size=&quot;size16&quot;&gt;docs.h2o.ai&lt;/p&gt;
&lt;/div&gt;
&lt;/a&gt;&lt;/figure&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;h2o 패키지의 장점은&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;1) 오픈소스이기 때문에 (코드 공개형) 무료로 사용 가능하다&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;2) 메모리 분산처리 등으로 빠르게 모델 학습이 가능하다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;3) &lt;span&gt;오버피팅 발생시 &lt;/span&gt;중지하는 옵션을 설정할 수 있어, 정확도가 조금 더 높은 모델을 얻을 수 있다.&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;4) 여러 언어로 활용 가능하다 (R, Python, 스칼라 등)&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;정도로 정리해볼 수 있을 것 같다. (+추가 장점이 있다면 댓글로 알려주시면 추가하겠습니다)&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;기본적으로 JAVA 언어로 작성되었지만 R, Python, JAVA 등의 언어로 모두 사용 가능하도록 구성되어 있기에 범용성이 높고, 언어별로 특성을 반영하여 코드명이 작성되어있다. 또한 Rest API를 지원하기에 json을 통해서 외부 프로그램 및 스크립트에서도 사용 가능하고 웹 ui를 지원하여 gui 기반으로 활용이 가능하다. (해당 부분은 따로 포스팅 할 예정)&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;하둡이나 스파크에서 작동 가능하기에 현업에서도 많이 사용하고 있다. (적어도 프로젝트를 수행한 기업에서는 실제로 쓰고 있었다)&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;기본적으로 맥이나 리눅스에서 사용하는 것을 권장한다. 윈도우에선 XGBoost 가 작동하지 않는 것 같다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;설치에 대한 포스팅은 따로 요청이 있을 경우 하려고 하며, 이번 포스팅에서는 데이터를 호출하여 xgboost 모델을 구축하는 방법에 대해 다루려고 한다. 함께 간단한 실습을 해보는 게 가장 좋을 거 같아서, colab 에서 간단한 예제와 함께 실행해보려 한다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;간단한 예제라면 역시 타이타닉 데이터다. XGBoost이기 때문에 categorical 변수가 있으면 좋을 거 같아서 선택했다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;설치&amp;nbsp;&lt;/h3&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;환경에 따라 다르지만, colab에서는&amp;nbsp;&lt;/p&gt;
&lt;pre id=&quot;code_1661928175646&quot; class=&quot;bash&quot; data-ke-language=&quot;bash&quot; data-ke-type=&quot;codeblock&quot;&gt;&lt;code&gt;!pip install h2o&lt;/code&gt;&lt;/pre&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;코드면 설치가 된다.&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;1.PNG&quot; data-origin-width=&quot;1831&quot; data-origin-height=&quot;457&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/YUcRy/btrK1hca818/QXm0UT9NNC0OJkU4BkqKYk/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/YUcRy/btrK1hca818/QXm0UT9NNC0OJkU4BkqKYk/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/YUcRy/btrK1hca818/QXm0UT9NNC0OJkU4BkqKYk/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FYUcRy%2FbtrK1hca818%2FQXm0UT9NNC0OJkU4BkqKYk%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;1831&quot; height=&quot;457&quot; data-filename=&quot;1.PNG&quot; data-origin-width=&quot;1831&quot; data-origin-height=&quot;457&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;local 컴퓨터에서 사용시 JAVA 프로그램(JVM)을 필수 사전 설치해야 한다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;환경별 자세한 설치 방법은 해당 페이지를 참고. (&lt;a href=&quot;https://docs.h2o.ai/h2o/latest-stable/h2o-docs/downloading.html&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot;&gt;https://docs.h2o.ai/h2o/latest-stable/h2o-docs/downloading.html&lt;/a&gt;)&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;실습에 필요한 패키지/데이터 호출 및 전처리&amp;nbsp;&lt;/h3&gt;
&lt;pre id=&quot;code_1661928361374&quot; class=&quot;python&quot; data-ke-language=&quot;python&quot; data-ke-type=&quot;codeblock&quot;&gt;&lt;code&gt;import pandas as pd 
import seaborn as sns 
import xgboost as xgb
import matplotlib.pyplot as plt
import h2o
from sklearn.model_selection import train_test_split&lt;/code&gt;&lt;/pre&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;사용 예정인 패키지를 모두 불러온다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;+ titanic 데이터는 seaborn 패키지 안에 내장되어 있다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;- 데이터 로드&amp;nbsp;&lt;/p&gt;
&lt;pre id=&quot;code_1661928452332&quot; class=&quot;python&quot; data-ke-language=&quot;python&quot; data-ke-type=&quot;codeblock&quot;&gt;&lt;code&gt;df = sns.load_dataset('titanic')
df&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;2.PNG&quot; data-origin-width=&quot;878&quot; data-origin-height=&quot;645&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/vtF2I/btrK49cTswu/zrvZe7pQ2zbWKK1QR0fkh1/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/vtF2I/btrK49cTswu/zrvZe7pQ2zbWKK1QR0fkh1/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/vtF2I/btrK49cTswu/zrvZe7pQ2zbWKK1QR0fkh1/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FvtF2I%2FbtrK49cTswu%2FzrvZe7pQ2zbWKK1QR0fkh1%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;878&quot; height=&quot;645&quot; data-filename=&quot;2.PNG&quot; data-origin-width=&quot;878&quot; data-origin-height=&quot;645&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;- 아까 말했듯 tree 기반이 아닌 다른 모델에서는 표준화, 결측값 처리 등의 절차가 추가되어야 하지만, 여기서는 사용하지 않는다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;- one-hot encoding 과 같은 가변수화는 h2o 에서 알아서 수행해준다. 너무 편한 기능이다.&amp;nbsp;&lt;/p&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;h2o 사용&lt;/h3&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;1) h2o 초기화&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;h2o 를 사용하려면 우선 초기화를 먼저 진행해야 한다. 아니면 오류가 발생할 수 있다.&amp;nbsp;&lt;/p&gt;
&lt;pre id=&quot;code_1661928682269&quot; class=&quot;python&quot; data-ke-language=&quot;python&quot; data-ke-type=&quot;codeblock&quot;&gt;&lt;code&gt;h2o.init()&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;4.PNG&quot; data-origin-width=&quot;877&quot; data-origin-height=&quot;668&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/rer6y/btrK1hca84k/jA3unqKEidDam82a2FyyUk/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/rer6y/btrK1hca84k/jA3unqKEidDam82a2FyyUk/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/rer6y/btrK1hca84k/jA3unqKEidDam82a2FyyUk/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2Frer6y%2FbtrK1hca84k%2FjA3unqKEidDam82a2FyyUk%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;877&quot; height=&quot;668&quot; data-filename=&quot;4.PNG&quot; data-origin-width=&quot;877&quot; data-origin-height=&quot;668&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;2) h2o 데이터 프레임화&amp;nbsp;&lt;/p&gt;
&lt;pre id=&quot;code_1661928870653&quot; class=&quot;python&quot; data-ke-language=&quot;python&quot; data-ke-type=&quot;codeblock&quot;&gt;&lt;code&gt;df_h2o = h2o.H2OFrame(df) #.import_file('titanic.csv')
df_h2o&lt;/code&gt;&lt;/pre&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;데이터를 h2o에서 사용할 수 있도록 H2OFrame화 시켜줘야 한다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;3) 변수형 변환&lt;/p&gt;
&lt;pre id=&quot;code_1661928896182&quot; class=&quot;python&quot; data-ke-language=&quot;python&quot; data-ke-type=&quot;codeblock&quot;&gt;&lt;code&gt;df_h2o[&quot;survived&quot;] = df_h2o[&quot;survived&quot;].asfactor()
df_h2o[['age', 'fare', 'sibsp']] = df_h2o[['age', 'fare', 'sibsp']].asnumeric()

# target 변수와 분리
x = df_h2o.columns
y = &quot;survived&quot;
x.remove(y)&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;5.PNG&quot; data-origin-width=&quot;869&quot; data-origin-height=&quot;504&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/bAt783/btrK1hXCDvq/ThIIK7DnS0lkzd0xGXK2j1/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/bAt783/btrK1hXCDvq/ThIIK7DnS0lkzd0xGXK2j1/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/bAt783/btrK1hXCDvq/ThIIK7DnS0lkzd0xGXK2j1/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FbAt783%2FbtrK1hXCDvq%2FThIIK7DnS0lkzd0xGXK2j1%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;869&quot; height=&quot;504&quot; data-filename=&quot;5.PNG&quot; data-origin-width=&quot;869&quot; data-origin-height=&quot;504&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;target 변수인 survived 변수는 factor변수기 때문에 먼저 바꿔주고, 컬럼명을 x에 넣고 target 변수이름을 지우는 remove를 수행한 결과를 보면 사용할 변수명들이 잘 모여있다.&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;4) train/valid/test 데이터 분할&lt;/p&gt;
&lt;pre id=&quot;code_1661929008556&quot; class=&quot;python&quot; data-ke-language=&quot;python&quot; data-ke-type=&quot;codeblock&quot;&gt;&lt;code&gt;train,test,valid = df_h2o.split_frame(ratios=[.7, .15],seed=1234)&lt;/code&gt;&lt;/pre&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;훈련, 검증, 테스트 데이터를 분할하는 기능도 h2o에 있다. split_frame 으로, 각각의 분할의 비율을 ratios 안에 적으면 비율에 맞게 분할 된다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;6.PNG&quot; data-origin-width=&quot;883&quot; data-origin-height=&quot;530&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/szbMW/btrKZ7H282A/0NUpx7PwbOk5l4tKoBWKCk/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/szbMW/btrKZ7H282A/0NUpx7PwbOk5l4tKoBWKCk/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/szbMW/btrKZ7H282A/0NUpx7PwbOk5l4tKoBWKCk/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FszbMW%2FbtrKZ7H282A%2F0NUpx7PwbOk5l4tKoBWKCk%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;883&quot; height=&quot;530&quot; data-filename=&quot;6.PNG&quot; data-origin-width=&quot;883&quot; data-origin-height=&quot;530&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;5) 기본적인 XGBoost 모델 구축&lt;/p&gt;
&lt;pre id=&quot;code_1661929354869&quot; class=&quot;python&quot; data-ke-language=&quot;python&quot; data-ke-type=&quot;codeblock&quot;&gt;&lt;code&gt;from h2o.estimators import H2OXGBoostEstimator

titanic_xgb = H2OXGBoostEstimator(seed=1234, 
                                  ntrees = 300, 
                                  max_depth = 5)&lt;/code&gt;&lt;/pre&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;h2o의 xgboost 구축 함수는 &lt;span style=&quot;color: #d4d4d4;&quot;&gt;H2OXGBoostEstimator &lt;/span&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;이다. 기본 hyper-parameter로 ntrees(tree의 최대개수), max_depth(tree의 깊이)가 있다.&amp;nbsp;&lt;/span&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;이 함수는 hyper-parameter를 설정하는 부분이고 실제 데이터로 모델을 훈련시키는 코드는 다음과 같다.&amp;nbsp;&lt;/span&gt;&lt;/p&gt;
&lt;pre id=&quot;code_1661929459001&quot; class=&quot;python&quot; data-ke-language=&quot;python&quot; data-ke-type=&quot;codeblock&quot;&gt;&lt;code&gt;titanic_xgb.train(x=x,
                  y=y,
                  training_frame=train,
                  validation_frame = valid)
                  # 추가 early stopping 옵션
                  #stopping_metric = &quot;auc&quot;, 
                  #stopping_tolerance=0.1, 
                  #stopping_rounds=3)&lt;/code&gt;&lt;/pre&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;아까 저장한 x에는 요인변수들이 들어있고, y에는 target 변수, training_frame에는 train에 활용할 데이터, validation_frame에는 validation 에 활용할 데이터를 넣어주면 된다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;7.PNG&quot; data-origin-width=&quot;849&quot; data-origin-height=&quot;773&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/bjCs3j/btrK1hQMTVH/Hp2Jw1RgpPOKUePM9VTE21/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/bjCs3j/btrK1hQMTVH/Hp2Jw1RgpPOKUePM9VTE21/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/bjCs3j/btrK1hQMTVH/Hp2Jw1RgpPOKUePM9VTE21/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FbjCs3j%2FbtrK1hQMTVH%2FHp2Jw1RgpPOKUePM9VTE21%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;849&quot; height=&quot;773&quot; data-filename=&quot;7.PNG&quot; data-origin-width=&quot;849&quot; data-origin-height=&quot;773&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;h2o의 장점이었던 early stopping 옵션은 모델마다 다르나, xgboost에서 가장 많이 사용하는 옵션은 위의 3가지로 stopping_metric 은 모델 훈련시 과적합 발생할 때 어떤 지표를 기준으로 종료하느냐를 설정하는 옵션이고, stopping_tolearance는 stopping round에서 이동평균이 계산되는데, 최상의 이동평균과 이 &lt;span&gt;&lt;span&gt;&amp;nbsp;&lt;/span&gt;stopping round 이동평균의 차이가 설정값 이상으로 벌어지는 경우 모델 훈련을 종료하는 옵션이다.&amp;nbsp;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;8.PNG&quot; data-origin-width=&quot;688&quot; data-origin-height=&quot;595&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/Sv3ky/btrK49jDZ4x/rUwEPFcx3s1bNKXhqa8uS0/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/Sv3ky/btrK49jDZ4x/rUwEPFcx3s1bNKXhqa8uS0/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/Sv3ky/btrK49jDZ4x/rUwEPFcx3s1bNKXhqa8uS0/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FSv3ky%2FbtrK49jDZ4x%2FrUwEPFcx3s1bNKXhqa8uS0%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;688&quot; height=&quot;595&quot; data-filename=&quot;8.PNG&quot; data-origin-width=&quot;688&quot; data-origin-height=&quot;595&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;모델 결과를 보면 굉장히 긴데, training과 validation 의 결과를 각각 보여주기 때문에 중복값으로 보이나 다르다.&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;9.PNG&quot; data-origin-width=&quot;796&quot; data-origin-height=&quot;582&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/bfttNz/btrK5E4L80V/y0RJog03dyCA0UiMzrBPt1/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/bfttNz/btrK5E4L80V/y0RJog03dyCA0UiMzrBPt1/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/bfttNz/btrK5E4L80V/y0RJog03dyCA0UiMzrBPt1/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FbfttNz%2FbtrK5E4L80V%2Fy0RJog03dyCA0UiMzrBPt1%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;796&quot; height=&quot;582&quot; data-filename=&quot;9.PNG&quot; data-origin-width=&quot;796&quot; data-origin-height=&quot;582&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;다양한 지표들을 확인할 수 있다.&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;tree 모형이기 때문에 나무의 구조에 따른 성능을 중간중간 보여주는 table도 출력된다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;10.PNG&quot; data-origin-width=&quot;1743&quot; data-origin-height=&quot;770&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/YRgBQ/btrKZ8GO56n/tkEtzWsRGsnsnroWjL8Du0/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/YRgBQ/btrKZ8GO56n/tkEtzWsRGsnsnroWjL8Du0/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/YRgBQ/btrKZ8GO56n/tkEtzWsRGsnsnroWjL8Du0/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FYRgBQ%2FbtrKZ8GO56n%2FtkEtzWsRGsnsnroWjL8Du0%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;1743&quot; height=&quot;770&quot; data-filename=&quot;10.PNG&quot; data-origin-width=&quot;1743&quot; data-origin-height=&quot;770&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;6) test&amp;nbsp;&lt;/p&gt;
&lt;pre id=&quot;code_1661932312198&quot; class=&quot;python&quot; data-ke-language=&quot;python&quot; data-ke-type=&quot;codeblock&quot;&gt;&lt;code&gt;pred = titanic_xgb.predict(test)
pred&lt;/code&gt;&lt;/pre&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;test데이터를 모델에 넣어서 예측하게 되면 class 가 나오는데, 이를 가지고 원래 test 데이터의 정답과 비교하면 성능을 얻을 수 있다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;11.PNG&quot; data-origin-width=&quot;892&quot; data-origin-height=&quot;403&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/XM4Jq/btrK4DMnmXh/k2K1LnoeZPa3aA4RzFcDX1/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/XM4Jq/btrK4DMnmXh/k2K1LnoeZPa3aA4RzFcDX1/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/XM4Jq/btrK4DMnmXh/k2K1LnoeZPa3aA4RzFcDX1/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FXM4Jq%2FbtrK4DMnmXh%2Fk2K1LnoeZPa3aA4RzFcDX1%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;892&quot; height=&quot;403&quot; data-filename=&quot;11.PNG&quot; data-origin-width=&quot;892&quot; data-origin-height=&quot;403&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;여기까지가 h2o 패키지를 활용하여 python 에서 xgboost 모델을 구축하는 과정이다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;</description>
      <category>데이터과학/데이터분석</category>
      <category>h2o</category>
      <category>h2o파이썬</category>
      <category>python</category>
      <category>pythonh2o</category>
      <category>xgboost파이썬</category>
      <category>파이썬xgb조기중지</category>
      <author>통경</author>
      <guid isPermaLink="true">https://stat-cbc.tistory.com/41</guid>
      <comments>https://stat-cbc.tistory.com/41#entry41comment</comments>
      <pubDate>Wed, 31 Aug 2022 16:55:42 +0900</pubDate>
    </item>
    <item>
      <title>[서평] 《Do it! 알고리즘 코딩 테스트 파이썬 편》</title>
      <link>https://stat-cbc.tistory.com/40</link>
      <description>&lt;p data-ke-size=&quot;size16&quot;&gt;오늘 리뷰해볼 도서는 《Do it! 알고리즘 코딩 테스트 파이썬 편》 이다.&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;사실 데이터 분석가 직무를 위한 취업을 하면서 코딩 테스트라는 부분을 준비할 생각은 하지 못했었다. 그러나 대학원을 졸업하고 다시 취업 준비를 하다보니 기본적으로 등장하는 부분이 코딩 테스트였다.&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;코딩테스트란 자료 구조 수업에 대한 선행지식이 없이 입문하는 나와 같은 데이터 분석가들에게는 엄청난 장벽같이 느껴진다.&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;특히, 파이썬으로 진입하게 되면 다른 언어와는 또 다른 문제가 발생한다(타 언어보다 긴 시간, pypy3 등) 이런 경우 사실 공부를 포기하게 되는 경우가 많은데, 포기하게 되면 대기업 중에 아예 코딩테스트 때문에 원서를 포기해야 하는 경우가 생긴다.&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;하지만 데이터 분석가로써, 모델러로써 빅데이터를 처리해야 하는 과정에서 코딩 테스트와 같이 효율성을 고려한 코드를 짤 필요가 있기에 코딩 테스트도 마냥 딴 세상 공부는 아니라고 생각한다.&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;데이터 분석가 직무에서는 파이썬이 가장 인기가 많은 언어라고 생각하는데, 본인이 기존에 R을 사용하다가 파이썬으로 전향한 경우이기 때문이다.&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;이 알고리즘 코딩 테스트 서적의 경우 데이터 분석 직무의 초심자가 보기에 적절한 책이라고 생각한다.&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;기존의 코딩테스트 서적 중 유명한 서적이&lt;/p&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;나동빈, 이것이 코딩테스트다&lt;/li&gt;
&lt;li&gt;권국원, 보통의 취준생을 위한 코딩 테스트 with 파이썬&lt;/li&gt;
&lt;/ul&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;개인적으로 이 두 권 정도를 봤는데, 코딩테스트다는 입문으로 쓰기에 아주 좋은 책이었으나 자료구조에 대한 자세한 설명이나 예제 문제가 조금 부족한 느낌을 받았었고 보통의 취준생을 위한&amp;hellip; 은 너무 자세하게 나와있고, 수학적인 내용이 많이 수록되어 있어 난이도가 굉장히 높다고 느껴졌다.&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;그러나, 해당 서적은 이 두 서적의 아쉬운 부분을 결합하여 조절해 둔 책이라고 느꼈다. 우선, 초심자들에게 일정을 짜는 테이블을 두어 루즈해지지 않도록하고, 각종 꿀팁들을 담아두어 좋았다.&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;KakaoTalk_20220829_222221549_01.jpg&quot; data-origin-width=&quot;1080&quot; data-origin-height=&quot;1440&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/cJ85Jd/btrKTPT2Auc/05Dvi1L7DUEYHDdnaFUoeK/img.jpg&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/cJ85Jd/btrKTPT2Auc/05Dvi1L7DUEYHDdnaFUoeK/img.jpg&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/cJ85Jd/btrKTPT2Auc/05Dvi1L7DUEYHDdnaFUoeK/img.jpg&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FcJ85Jd%2FbtrKTPT2Auc%2F05Dvi1L7DUEYHDdnaFUoeK%2Fimg.jpg&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;1080&quot; height=&quot;1440&quot; data-filename=&quot;KakaoTalk_20220829_222221549_01.jpg&quot; data-origin-width=&quot;1080&quot; data-origin-height=&quot;1440&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;사실 코딩테스트는 꾸준히 준비해야 하지만, 가장 효율이 좋을 때가 코테가 임박했을 때라고 생각한다. 그럴 때 사용 가능한 루틴까지 꼼꼼하게 정리해두어서 좋았다!&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;KakaoTalk_20220829_222221549_02.jpg&quot; data-origin-width=&quot;1080&quot; data-origin-height=&quot;1440&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/sECOr/btrKSgLd7fF/aQe9T1Ze9Ki821yDWquM7K/img.jpg&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/sECOr/btrKSgLd7fF/aQe9T1Ze9Ki821yDWquM7K/img.jpg&quot; data-alt=&quot;실제로 시험 전 이 루틴을 따라해보려 한다.&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/sECOr/btrKSgLd7fF/aQe9T1Ze9Ki821yDWquM7K/img.jpg&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FsECOr%2FbtrKSgLd7fF%2FaQe9T1Ze9Ki821yDWquM7K%2Fimg.jpg&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;1080&quot; height=&quot;1440&quot; data-filename=&quot;KakaoTalk_20220829_222221549_02.jpg&quot; data-origin-width=&quot;1080&quot; data-origin-height=&quot;1440&quot;/&gt;&lt;/span&gt;&lt;figcaption&gt;실제로 시험 전 이 루틴을 따라해보려 한다.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;사실 코딩테스트는 정답이 없다고 말할 수 있을 정도로 다양한 방식으로 코드가 작성되는데, 기본적인 원리인 시간복잡도와 디버깅에 대해 짚고 넘어가주어 깔끔했다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;KakaoTalk_20220829_222221549_03.jpg&quot; data-origin-width=&quot;1080&quot; data-origin-height=&quot;1440&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/cJCd5T/btrKUdUyYat/uMLICnNxn0YL1enzZkQQnK/img.jpg&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/cJCd5T/btrKUdUyYat/uMLICnNxn0YL1enzZkQQnK/img.jpg&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/cJCd5T/btrKUdUyYat/uMLICnNxn0YL1enzZkQQnK/img.jpg&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FcJCd5T%2FbtrKUdUyYat%2FuMLICnNxn0YL1enzZkQQnK%2Fimg.jpg&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;1080&quot; height=&quot;1440&quot; data-filename=&quot;KakaoTalk_20220829_222221549_03.jpg&quot; data-origin-width=&quot;1080&quot; data-origin-height=&quot;1440&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;코딩에서 제일 중요하다고 생각하는 디버깅 부분.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;KakaoTalk_20220829_222221549_04.jpg&quot; data-origin-width=&quot;1080&quot; data-origin-height=&quot;1440&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/Kf3Lz/btrKTPGvj6j/PVtIi68yGbvNUva7tN0jP1/img.jpg&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/Kf3Lz/btrKTPGvj6j/PVtIi68yGbvNUva7tN0jP1/img.jpg&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/Kf3Lz/btrKTPGvj6j/PVtIi68yGbvNUva7tN0jP1/img.jpg&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FKf3Lz%2FbtrKTPGvj6j%2FPVtIi68yGbvNUva7tN0jP1%2Fimg.jpg&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;1080&quot; height=&quot;1440&quot; data-filename=&quot;KakaoTalk_20220829_222221549_04.jpg&quot; data-origin-width=&quot;1080&quot; data-origin-height=&quot;1440&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;원리를 알고 배우니 훨씬 더 좋았다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;+ 깨알 빈출, 핵심 라벨도 달아주어 어떤 유형인지에 대해 알 수 있었다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-filename=&quot;edited_KakaoTalk_20220829_222221549_06.jpg&quot; data-origin-width=&quot;1080&quot; data-origin-height=&quot;1440&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/dfUvjD/btrKVWLidBY/EQ0OLNvMEBCSErhVJBcubk/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/dfUvjD/btrKVWLidBY/EQ0OLNvMEBCSErhVJBcubk/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/dfUvjD/btrKVWLidBY/EQ0OLNvMEBCSErhVJBcubk/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FdfUvjD%2FbtrKVWLidBY%2FEQ0OLNvMEBCSErhVJBcubk%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;1080&quot; height=&quot;1440&quot; data-filename=&quot;edited_KakaoTalk_20220829_222221549_06.jpg&quot; data-origin-width=&quot;1080&quot; data-origin-height=&quot;1440&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;총평으로는 코딩테스트 입문서로는 아주 적절한 것 같다고 생각한다. 단언컨데 현존하는 파이썬 코딩테스트 대비 서적 중 가장 이론과 실전 문제의 밸런스를 잘 맞춘 책이라고 할 수 있을 것 같다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;실제 이 책으로 코딩테스트 준비도 해보고 문제도 풀어보면서 차차 추가 후기를 남기도록 해야겠다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;background-color: #ffffff; color: #666666;&quot;&gt;​&lt;/span&gt;&lt;span style=&quot;background-color: #ffffff; color: #ffef00;&quot;&gt;※ 출판사의 &lt;/span&gt;&lt;span style=&quot;background-color: #ffffff; color: #ffef00;&quot;&gt;서평 이벤트로 책을 받아 서평을 작성하였습니다. 또한 , 이 글은 저의 주관적인 생각이 담겼습니다.&lt;/span&gt;&lt;/p&gt;</description>
      <category>취업준비/코딩테스트</category>
      <category>서평</category>
      <author>통경</author>
      <guid isPermaLink="true">https://stat-cbc.tistory.com/40</guid>
      <comments>https://stat-cbc.tistory.com/40#entry40comment</comments>
      <pubDate>Mon, 29 Aug 2022 22:38:26 +0900</pubDate>
    </item>
    <item>
      <title>[2019 카카오] 오픈채팅방</title>
      <link>https://stat-cbc.tistory.com/39</link>
      <description>&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;a href=&quot;https://programmers.co.kr/learn/courses/30/lessons/42888&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot;&gt;https://programmers.co.kr/learn/courses/30/lessons/42888&lt;/a&gt;&lt;/p&gt;
&lt;figure id=&quot;og_1654525244463&quot; contenteditable=&quot;false&quot; data-ke-type=&quot;opengraph&quot; data-ke-align=&quot;alignCenter&quot; data-og-type=&quot;website&quot; data-og-title=&quot;코딩테스트 연습 - 오픈채팅방&quot; data-og-description=&quot;오픈채팅방 카카오톡 오픈채팅방에서는 친구가 아닌 사람들과 대화를 할 수 있는데, 본래 닉네임이 아닌 가상의 닉네임을 사용하여 채팅방에 들어갈 수 있다. 신입사원인 김크루는 카카오톡 오&quot; data-og-host=&quot;programmers.co.kr&quot; data-og-source-url=&quot;https://programmers.co.kr/learn/courses/30/lessons/42888&quot; data-og-url=&quot;https://programmers.co.kr/learn/courses/30/lessons/42888&quot; data-og-image=&quot;https://scrap.kakaocdn.net/dn/TDQbH/hyOFAR7Ybb/5Kn0R76ydyfbxkqdwOS7K1/img.jpg?width=626&amp;amp;height=626&amp;amp;face=0_0_626_626,https://scrap.kakaocdn.net/dn/LeXS6/hyOFzeAnS5/kT4atc9YkgkbHWENmG6Bm1/img.jpg?width=626&amp;amp;height=626&amp;amp;face=0_0_626_626&quot;&gt;&lt;a href=&quot;https://programmers.co.kr/learn/courses/30/lessons/42888&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; data-source-url=&quot;https://programmers.co.kr/learn/courses/30/lessons/42888&quot;&gt;
&lt;div class=&quot;og-image&quot; style=&quot;background-image: url('https://scrap.kakaocdn.net/dn/TDQbH/hyOFAR7Ybb/5Kn0R76ydyfbxkqdwOS7K1/img.jpg?width=626&amp;amp;height=626&amp;amp;face=0_0_626_626,https://scrap.kakaocdn.net/dn/LeXS6/hyOFzeAnS5/kT4atc9YkgkbHWENmG6Bm1/img.jpg?width=626&amp;amp;height=626&amp;amp;face=0_0_626_626');&quot;&gt;&amp;nbsp;&lt;/div&gt;
&lt;div class=&quot;og-text&quot;&gt;
&lt;p class=&quot;og-title&quot; data-ke-size=&quot;size16&quot;&gt;코딩테스트 연습 - 오픈채팅방&lt;/p&gt;
&lt;p class=&quot;og-desc&quot; data-ke-size=&quot;size16&quot;&gt;오픈채팅방 카카오톡 오픈채팅방에서는 친구가 아닌 사람들과 대화를 할 수 있는데, 본래 닉네임이 아닌 가상의 닉네임을 사용하여 채팅방에 들어갈 수 있다. 신입사원인 김크루는 카카오톡 오&lt;/p&gt;
&lt;p class=&quot;og-host&quot; data-ke-size=&quot;size16&quot;&gt;programmers.co.kr&lt;/p&gt;
&lt;/div&gt;
&lt;/a&gt;&lt;/figure&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;카카오의 2019 블라인드의 오픈채팅방 문제이다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;입력 : 상호작용(Enter Leave Change) 각 유저의 아이디 닉네임&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;출력 : &quot;닉네임&quot; 님이 &quot;상호작용&quot; 하였습니다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;[문제풀이]&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;1. 단순히 문제만 잘 이해한다면 조건문으로 해결 가능한 문제&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;pre id=&quot;code_1654526362495&quot; class=&quot;python&quot; data-ke-language=&quot;python&quot; data-ke-type=&quot;codeblock&quot;&gt;&lt;code&gt;# record input : 명령어, 유저아이디, 닉네임 한 행 

def solution(record):
    answer = []
    trace = []
    map = {}
    for i in range(len(record)):
        temp = record[i].split(' ')
        if temp[0] == 'Enter':
            map[temp[1]] = temp[2]
            trace.append([temp[0], temp[1]])
        elif temp[0] == 'Leave':
            trace.append([temp[0], temp[1]])
        else:
            map[temp[1]] = temp[2]

    for i in range(len(trace)):
        if trace[i][0] == 'Enter':
            result = map[trace[i][1]] + &quot;님이 들어왔습니다.&quot;
            answer.append(result)
        else:
            result = map[trace[i][1]] + &quot;님이 나갔습니다.&quot;
            answer.append(result)

    return answer&lt;/code&gt;&lt;/pre&gt;</description>
      <category>취업준비/코딩테스트</category>
      <category>카카오블라인드</category>
      <category>카카오채용</category>
      <category>카카오코테</category>
      <category>카카오파이썬</category>
      <category>코딩테스트</category>
      <category>파이썬</category>
      <category>파이썬코테</category>
      <category>프로그래머스</category>
      <author>통경</author>
      <guid isPermaLink="true">https://stat-cbc.tistory.com/39</guid>
      <comments>https://stat-cbc.tistory.com/39#entry39comment</comments>
      <pubDate>Mon, 6 Jun 2022 23:39:52 +0900</pubDate>
    </item>
    <item>
      <title>[백준 15686] 치킨배달 Python 코드</title>
      <link>https://stat-cbc.tistory.com/38</link>
      <description>&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;a href=&quot;https://www.acmicpc.net/problem/15686&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot;&gt;https://www.acmicpc.net/problem/15686&lt;/a&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;백준의 15686 치킨배달은 삼성 SW역량테스트 A형 기출문제로 출제되었다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;문제는 N * N 크기의 도시에서 빈칸, 치킨집, 집 중 하나로 구성되어 있을 때, 집과 치킨집의 거리를 계산하여 도시에 있는 치킨집 중 치킨 거리가 가장 작게 될 지 구하는 프로그램을 작성하는 문제이다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;입력 : 첫째 줄에 N(2 &amp;lt;= N &amp;lt;= 50)과 M(1&amp;lt;= M&amp;lt;=13) 이 주어진다. 둘째 줄 부터 N개의 줄에는 도시 정보가 주어진다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;도시정보&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;- 0 : 빈칸&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;- 1 : 집&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;- 2 : 치킨집&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;출력 : 첫째 줄에 폐업시키지 않을 치킨집 최대 M개를 골랐을 때, 도시의 치킨 거리의 최솟값을 출력&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;[문제풀이]&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;1. 아직 코딩테스트 입문자라 처음 지도를 보고서는 구현 종류의 문제인줄 알았으나, 완전 탐색 문제(재귀함수)인 것을 생각해냈다.&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;2.&amp;nbsp; 우선 치킨 집과 집 간의 거리를 계산해야 하는 함수와 치킨 집을 고른 개수가 m일 때 조건을 주어 모든 집과 치킨 집 사이의 최단 거리를 구한 후 반환하는 함수 두 개를 구현해야 한다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;pre id=&quot;code_1654525951381&quot; class=&quot;python&quot; data-ke-language=&quot;python&quot; data-ke-type=&quot;codeblock&quot;&gt;&lt;code&gt;# 도시의 크기, 치킨 집 개수를 입력받음 
n,m = map(int, input().split())
# 도시 구조에 대한 정보를 입력받음 
City = [list(map(int, input().split())) for _ in range(n)]

def Brute_Force(idx, x, y): 
    global answer
    if(idx == m):
        choice_chicken = []
        for i in range(n):
            for j in range(n):
                if(City[i][j] == 3):
                    choice_chicken.append((i,j))
        res = Min_Distance(choice_chicken, house)
        if(answer &amp;gt; sum(res)):
            answer = sum(res)
        return 
    else:
        for i in range(x, n):
            if(i == x): k = y 
            else: k = 0
            for j in range(k,n):
                if(City[i][j] == 2):
                    City[i][j] = 3
                    Brute_Force(idx+1, i, j+1)
                    City[i][j] = 2

def Min_Distance(chicken, house):
    sum_Distance = []
    for i in house:
        min_D = 987654321
        for j in chicken:
            Distance = abs(i[0] - j[0]) + abs(i[1] - j[1])
            min_D = min(min_D, Distance)
        sum_Distance.append(min_D)
    return sum_Distance



answer = 987654321
house = []
for i in range(n):
    for j in range(n):
        if(City[i][j] == 1):
            house.append((i,j))
Brute_Force(0,0,0)
print(answer)&lt;/code&gt;&lt;/pre&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;</description>
      <category>취업준비/코딩테스트</category>
      <category>백준15686</category>
      <category>백준치킨거리</category>
      <category>삼성SW기출</category>
      <category>코딩테스트</category>
      <category>파이썬 코딩테스트</category>
      <category>파이썬코딩</category>
      <author>통경</author>
      <guid isPermaLink="true">https://stat-cbc.tistory.com/38</guid>
      <comments>https://stat-cbc.tistory.com/38#entry38comment</comments>
      <pubDate>Mon, 6 Jun 2022 23:33:10 +0900</pubDate>
    </item>
    <item>
      <title>Swin Transformer: Hierarchical Vision Transformer using Shifted Windows</title>
      <link>https://stat-cbc.tistory.com/37</link>
      <description>&lt;article id=&quot;81980ead-c7c1-490b-8e25-d74fc437f2f7&quot; class=&quot;page sans Notion_P&quot;&gt;
&lt;div class=&quot;page-body&quot;&gt;&lt;nav id=&quot;2b8c376f-347b-460d-910c-7ef9de7ca48c&quot; class=&quot;block-color-gray table_of_contents&quot;&gt;
&lt;div class=&quot;table_of_contents-item table_of_contents-indent-1&quot;&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;keypoint : computation complexity, shifted window&lt;/span&gt;&lt;/div&gt;
&lt;/nav&gt;
&lt;p id=&quot;0ad35f66-65c2-480f-b9e3-c7cdf17f4e52&quot; class=&quot;&quot; data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h2 id=&quot;60e2cebb-b5d5-4bc0-a62b-2ba960ad578e&quot; class=&quot;&quot; data-ke-size=&quot;size26&quot;&gt;Abstract&lt;/h2&gt;
&lt;ul id=&quot;171141ef-6f93-4cae-9a9c-300d6ca6ef0d&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;컴퓨터 비전의 범용 백본 역할을 할 수 있는 새로운 비전 트랜스포머(Swin Transformer)를 소개.
&lt;ul id=&quot;3d719bcf-488a-4d68-996a-966a9919ef5a&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;NLP에서 vision으로 트랜스포머를 적응(adapting)시키는 문제는 두 domains 간의 차이에서 발생.
&lt;ul id=&quot;6f94c7dd-1d66-4b36-bd36-3370c5dac4d6&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;예) visual entities의 scale의 visual entities와 텍스트의 단어와 비교하여 이미지의 픽셀 해상도가 높음.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;db11ba17-aa57-4406-898d-0d79362567be&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;이러한 차이를 해결하기 위해 논문은 representation이 &lt;b&gt;shifted windows&lt;/b&gt;으로 계산되는 &lt;b&gt;계층적 트랜스포머&lt;/b&gt;를 제안.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;22f64671-3bd9-496e-a200-53899a993bbf&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;shifted windowing scheme는 &lt;b&gt;cross-window connection을 허용&lt;/b&gt;하는 동시에 self-attention 계산을 &lt;b&gt;non-overlapping local windows으로 제한함으로써 효율성&lt;/b&gt; 향상.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;c85346e8-374b-4b2b-b156-7198cca0a1ad&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;이 계층 구조 아키텍처는 다양한 scales로 모델링할 수 있는 &lt;b&gt;flexibility&lt;/b&gt;과 image size에 대한 linear computational &lt;b&gt;complexity&lt;/b&gt;을 가짐.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;da5ff1c4-d472-4d20-a287-48fdea77a7ff&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;Swin Transformer의 성능은 이미지 분류(ImageNet-1K에서 86.4 top-1 정확도)와 object detection(58.7 box AP 및 51.1 mask AP on COCO test-dev) 및 semantic segmentation(53.5 mioU on ADEK on ADME 20)과 같은 dense prediction task을 포함한 광범위한 비전 과제와 호환가능&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;ddfd5802-f634-4300-99bb-9da8554d0e7e&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;성능은 COCO에서는 +2.7 박스 AP 및 +2.6 마스크 AP, ADE20K에서는 +3.2 mIoU의 큰 차이로 이전 SOTA 모델을 능가하며 vision 백본으로서의 트랜스포머 기반 모델의 잠재력을 보여줌.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;b389d771-d0fa-45ce-b590-b9c8a909acf5&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;코드와 모델은 https://github.com/microsoft/Swin-Transformer에서 공개.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;d2a32355-51ec-442c-a8b0-0d4bea5c92e1&quot; class=&quot;&quot; data-ke-size=&quot;size26&quot;&gt;Introduction&lt;/h2&gt;
&lt;ul id=&quot;d23c9b35-f9f7-463b-befd-229624e8bb3b&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;computer vision에서의 모델링은 오랫동안 CNN에 의해 지배되어옴.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;ba9f7ee7-5f0a-4bf4-8647-15eb9b01e0aa&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;ImageNet 이미지 분류 문제에 대한 AlexNet [38]과 혁신적인 성능을 시작으로, CNN 아키텍처는 greater scale[29, 73], more extensive connections[33], more extensive connections[67, 17, 81]을 통해 점점 더 강력해짐.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;736cd30c-1602-4a2c-94e7-176e9cb1ca5d&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;CNN이 다양한 vision tasks를 위한 백본 네트워크 역할을 하는 가운데, 이러한 아키텍처의 발전은 전체 분야를 광범위하게 끌어올린 성능 향상으로 이어짐.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;d16c4e2b-107f-47f3-b735-12bd902ba7d9&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;반면에, 자연 언어 처리(NLP)에서의 네트워크 아키텍처의 진화는 다른 경로를 따름.
&lt;ul id=&quot;9b081034-6fe3-4f00-8ff1-33b5fbbf76a6&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;보편적인 아키텍처 : Transformer[61].&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;9c7fe4a1-18f9-4978-9a73-bf0a912af5f5&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;시퀀스 모델링 및 변환 작업을 위해 설계된 Transformer는 장거리 의존성 모델링에 attention 사용.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;856738ef-0322-4848-a27f-1eecfe980460&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;최근 image classification[19]와 joint vision-language modeling[46]과 같은 유망한 결과를 보여줌.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;c163ff4d-8975-43c5-b52c-c3816b6bf1a3&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;본 논문은 Transformer 가 NLP와 vision에서 범용 백본 역할을 할 수 있도록 함.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;aa2d08f6-7918-463a-acbb-6d83faca366a&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;언어 영역의 고성능을 시각적 영역으로 전달하는 데 있어 중요한 문제는 두 &lt;b&gt;modalities (image &amp;harr; text) &lt;/b&gt;간의 차이로 볼 수 있음.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;fcde9952-e415-4e64-9d0c-bfbf59797069&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;차이점
&lt;ul id=&quot;751ffe49-4200-492f-9f9a-45bda7875b26&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;1) scale
&lt;ul id=&quot;c1173901-a901-4ae3-b987-46c0ab86a05b&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;language Transformer 처리의 기본 요소 역할을 하는 단어 토큰과는 달리, object detection[41, 52, 53]의 시각적 요소는 &lt;b&gt;scale&lt;/b&gt;에서 달라질 수 있음.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;4401e521-98d9-4d91-9b8c-6ebe7b3d8a2d&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;기존의 Transformer 기반 모델[61, 19]에서 &lt;b&gt;토큰&lt;/b&gt;은 모두 고정된 &lt;b&gt;scale이므로&lt;/b&gt; vision applications에 적합하지 않음.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;d1f6776b-ee6a-4263-aac3-a5ef423a7698&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;2) 해상도
&lt;ul id=&quot;0b80cf8a-3125-4a6e-9dc4-72899b1013ca&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;passages of text의 words에 비해 이미지의 픽셀 해상도가 훨씬 높음.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;3adeb7d4-1589-4ca2-adde-e9bc070e4f54&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;semantic segmentation : 픽셀 수준에서 dense prediction 이 필요 &amp;gt; self-attention의 계산 복잡성이 이미지 크기에 &lt;mark class=&quot;highlight-orange&quot;&gt;&lt;b&gt;quadratic&lt;/b&gt;&lt;/mark&gt;&lt;b&gt;계산 복잡도&lt;/b&gt;를 갖기 때문에 고해상도 이미지에서는 Transformer가 다루기 어려움.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;33c9c280-c388-41c9-b8dd-9897ba4991e4&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;이러한 문제를 해결하기 위해 논문은 &lt;b&gt;hierarchical feature maps&lt;/b&gt;을 구성하고 이미지 크기에 &lt;mark class=&quot;highlight-orange&quot;&gt;&lt;b&gt;Linear &lt;/b&gt;&lt;/mark&gt;&lt;b&gt;계산 복잡도&lt;/b&gt;를 갖는 범용 Transformer backbone인 Swin Transformer를 제안.
&lt;figure id=&quot;e906f757-861c-45be-a3c7-df74ce24c58c&quot; class=&quot;image&quot;&gt;&lt;a href=&quot;https://i.imgur.com/zrQ46ny.png[/img]&quot;&gt;&lt;img style=&quot;width: 588px;&quot; src=&quot;https://i.imgur.com/zrQ46ny.png[/img]&quot; /&gt;&lt;/a&gt;
&lt;figcaption&gt;그림 1. (a) 제안된 Swin Transformer는 이미지 패치(회색)를 더 깊은 계층으로 병합하여 계층적 피 맵을 구축하며, 각 local window(빨간색) 내에서만 self-attention의 계산으로 인해 이미지 크기를 입력하기 위한 linear computation complexity을 가짐. 따라서 이미지 분류와 dense recognition tasks 모두에 범용 백본 역할을 할 수 있음. (b) 반대로, 이전 비전 Transformers[19]는 feature maps of a single low resolution을 생성하며, 전체적으로 self-attention 의 계산으로 인해 영상 크기를 입력하기 위한 quadratic computation complexity을 가짐.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;ul id=&quot;fc73988a-190c-4888-9272-6dec86c8a32e&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;+ Vit 사진을 동일한 걸 쓴 이유?&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;35e5904d-fab1-4789-9960-aff737cc0242&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;그림 1(a)에서 볼 수 있듯이, Swin Transformer는 작은 크기의 patch(회색)에서 시작하여 점차적으로 더 상위(deeper) 트랜스포머 계층에서 인접 patch를 병합하여 계층적 feature를 구성.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;ac3ee7e6-fd75-4c45-ba16-a1844dd6817e&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;이러한 계층적 feature map을 통해, Swin Transformer 모델은 feature pyramid networks (FPN) [41] 또는&lt;b&gt; U-Net&lt;/b&gt; [50]과 같은 dense prediction을 위한 고급 기술을 편리하게 활용할 수 있음.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;9201e1d8-0c89-4ed5-8f4a-62ef2725fba9&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;&lt;b&gt;linear &lt;/b&gt;computational complexity은 이미지를 분할하는 &lt;b&gt;non-overlapping windows&lt;/b&gt;(빨간색) 내에서 locally 하게 self-attention를 계산함으로써 달성.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;f9b0f805-0220-4455-9d46-d6ee1e28eeb3&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;&lt;mark class=&quot;highlight-orange&quot;&gt;각 windows의 패치 수가 고정되므로 complexity가 &lt;/mark&gt;&lt;mark class=&quot;highlight-orange&quot;&gt;&lt;b&gt;이미지 크기에 비례&lt;/b&gt;&lt;/mark&gt;&lt;mark class=&quot;highlight-orange&quot;&gt;. &lt;/mark&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;c91df937-652b-40a7-9441-8cf5184fc031&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;이러한 장점 때문에 Swin Transformer는 single resolution의 feature map을 생성하며&lt;b&gt; quadratic complexity&lt;/b&gt;를 갖는 이전의 Transformer 기반 아키텍처[19]와는 달리 다양한 비전 작업에 대한 범용 백본으로서 적합.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;7e4b03d7-5063-424e-bee0-d4ac8f47a88a&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Swin Transformer의 핵심 설계 요소는 그림 2에서와 같이 연속적인 self-attention 계층 간 window partition의 이동.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;bae58c04-3922-4f96-a32d-4eb53e4e002a&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;shifted window 은 이전 레이어의 windows를 브리지하여 모델링 power를 크게 향상시키는 연결을 제공(Table 4 참조).&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;b0047207-ff1f-4952-95da-b519fe7af597&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;이 전략은 실제 대기 시간과 관련해서도 효율적.
&lt;ul id=&quot;8ee2ee1b-e572-4fd9-aec5-c7d37498525a&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;즉, window 내의 모든 query patch는 동일한 key sets (The query and key are projection vectors in a self-attention layer)를 공유하므로 하드웨어의 메모리 액세스가 용이해짐.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;d4d9c35d-13b7-4c2c-a34f-fd4c498c4139&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;반면, 이전의 sliding window 기반 self-attention approaches[32, 49]은 서로 다른 query pixel에 대해 서로 다른 key sets로 인해 일반 하드웨어에서 짧은 지연 시간이 발생함.
&lt;ul id=&quot;e89e6c6f-3ecc-4512-a692-1536cbf34294&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;논문의 실험을 통해 제안된 shifted window 접근 방식은 sliding window 방식보다 지연 시간이 훨씬 짧지만 모델링 power은 비슷하다는 것을 알 수 있음(표 5와 6 참조).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;f48f4056-982e-4aac-853d-d71ae85c0ad7&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;
&lt;ul id=&quot;15981c0c-44f6-4698-8fe6-32f970018c65&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;ADE20K semantic segmentation의 경우, val set에서 53.5mIoU를 얻는데, 이는 이전 state-of-the-art(SETR [78])에 비해 +3.2mIoU가 개선됨.제안된 Swin Transformer는 이미지 분류, 객체 감지 및 의미 분할의 인식 작업에서 강력한 성능을 달성.
&lt;ul id=&quot;7512d416-c28e-4265-8277-50e632314e19&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;ViT/DeT[19, 60] 및 ResNe(X)t 모델[29, 67]보다 성능이 뛰어나며 세 가지 작업에서 유사한 지연 시간이 발생.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;029355ad-3e31-4e82-87e2-0a15edad6e50&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;COCO test-dev set의 58.7 box AP 및 51.1 mask AP는 +2.7 box AP(외부 데이터가 없는 Copy-paste [25]) 및 +2.6 mask AP(DetectorRS [45])로 이전 SOTA 결과를 능가.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;54433a54-19c3-4fd7-98a2-13f420496b8f&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;ImageNet-1K 이미지 분류에서 86.4%의 정확도를 달성.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;9f1f7aa2-e42a-4310-95f8-65ab8d2da336&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;컴퓨터 비전과 자연 언어 처리에 걸친 통합 아키텍처는 시각 신호와 텍스트 신호의 공동 모델링을 촉진하고 두 도메인의 모델링 지식을 더 깊이 공유할 수 있기 때문에 양쪽 모두에 도움이 될 수 있다고 생각함.
&lt;ul id=&quot;cdb07c60-ec72-4618-87ec-70ef51eac7cc&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;논문은 다양한 비전 문제에 대한 Swin Transformer의 강력한 성능이 커뮤니티에 이러한 믿음을 더 깊이 심어주고 비전 및 언어 신호의 통합 모델링을 장려할 수 있기를 바람.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;4a83fa7e-35ca-4eb8-a431-bcbebb14b265&quot; class=&quot;&quot; data-ke-size=&quot;size26&quot;&gt;&amp;nbsp;Related Work&lt;/h2&gt;
&lt;ul id=&quot;6866990a-779c-4152-84f3-6f8009f4dc2e&quot; class=&quot;toggle&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;&lt;details&gt;
&lt;summary&gt;CNN and variants&lt;/summary&gt;
&lt;ul id=&quot;983cd53f-4f50-4081-9985-10cf78e11cf0&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;CNN은 컴퓨터 비전 전반에 걸쳐 표준 네트워크 모델 역할. CNN이 수십 년 동안 존재했지만 [39] AlexNet을 도입하고 나서야 CNN이 주류가 됨. 그 이후, 컴퓨터 비전의 딥러닝 파동을 더욱 촉진하기 위해 deeper and more effective convolutional neural architectures가 제안됨. VGG [51], GoogleNet [56], ResNet [29], DenseNet [33],HRNet [62] 및 EfficientNet [57].&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;c34da2a8-1faf-4aad-a3b5-7f46ce434605&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;이러한 아키텍처의 진보 외에도, depthwise convolution [67] 및 deformable convolution [17, 81]과 같은 개별 컨볼루션 레이어의 개선에 대한 많은 연구가 있었음.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;545923d1-8bcc-4ab3-93b6-602249f43680&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;CNN과 그 variants는 여전히 컴퓨터 비전 애플리케이션의 주요 백본 아키텍처이지만, 우리는 시각과 언어 사이의 통합 모델링을 위한 트랜스포머와 같은 아키텍처의 강력한 잠재력을 강조. 저자는 이 작업이 몇 가지 기본적인 시각적 인식 작업에서 강력한 성과를 달성하며 모델링 전환에 기여하기를 바람.&lt;/li&gt;
&lt;/ul&gt;
&lt;/details&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;636ebc33-c376-4dd7-b957-0dc2036fe28f&quot; class=&quot;toggle&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;&lt;details&gt;
&lt;summary&gt;Self-attention based backbone architectures&lt;/summary&gt;
&lt;ul id=&quot;75172408-f552-4473-b7c0-a5dbc8746850&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;또한 NLP 분야에서 Self-attention layers 과 Transformer architectures의 성공에 영감을 받아, 일부 작업은 인기 있는 ResNet에서 공간적 변환 계층의 일부 또는 전부를 대체하기 위해 Self-attention layers을 사용[32, 49, 77].&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;75e6fdbb-9574-4aa8-a537-59816ac1b914&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;이 작업에서는 각 픽셀의 local windows 내에서 Self-attention 를 계산하여 최적화[32]를 촉진하고, counterpart ResNet 아키텍처보다 약간 더 나은 accuracy/FLOPs trade-offs를 달성. 그러나, costly memory access는 실제 대기 시간이 convolutional networks의 대기 시간보다 훨씬 더 커지게 함[32].&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;f8b1fbe5-06c1-443c-b85a-3bb9c9ff3450&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;sliding windows를 사용하는 대신 consecutive layers 간에 shift windows를 사용하여 일반 하드웨어에서 보다 효율적으로 구현할 수 있도록 제안.&lt;/li&gt;
&lt;/ul&gt;
&lt;/details&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;cb1d76c0-95b2-431b-bc76-1ac0eb4201fb&quot; class=&quot;toggle&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;&lt;details&gt;
&lt;summary&gt;Self-attention/Transformers to complement CNNs&lt;/summary&gt;
&lt;ul id=&quot;11559577-84b7-4acc-a33d-cd6073069a55&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;표준 CNN 아키텍처를 self-attention layers 또는 Transformers로 강화하는 것. self-attention layers은 distant dependencies 또는 heterogeneous interactions을 인코딩하는 기능을 제공하여 complement backbones[64, 6, 68, 22, 71, 54] 또는 head networks[31, 26]를 보완할 수 있음. 보다 최근에는 트랜스포머의 encoder-decoder 설계가 object detection and instance segmentation tasks에 적용[7, 12, 82, 55]. 본 연구에서는 기본적인 visual feature extraction을 위한 트랜스포머의 adaptation을 살펴보고 이러한 작업을 보완함.&lt;/li&gt;
&lt;/ul&gt;
&lt;/details&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;18776313-8c2a-4709-83b9-1344a3b7b56e&quot; class=&quot;toggle&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;&lt;details&gt;
&lt;summary&gt;Transformer based vision backbones&lt;/summary&gt;
&lt;ul id=&quot;867f9684-24c3-4dc7-9933-7659117dc94d&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;ViT(Vision Transformer)[19]와 그 후속연구[60, 69, 14, 27, 63]와 관련.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;a301f22e-a66f-4f28-949e-a801c8e2544c&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;ViT의 pioneering work는 이미지 분류를 위해 겹치지 않는 중간 크기의 이미지 패치에 Transformer 아키텍처를 직접 적용. Convolutional 네트워크와 비교하여 이미지 분류에서 놀라운 speed-accuracy trade-off를 이룸. ViT는 우수한 성능을 발휘하려면 대규모 training datasets(즉, JFT-300M)이 필요하지만 DeiT[60]는 더 작은 ImageNet-1K 데이터셋을 사용하여 ViT를 효과적으로 운영할 수 있는 여러 training strategies을 도입.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;60934b25-0079-492c-aaa8-92bd66eaf445&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;ViT image classification 결과는 encouraging하지만, 이 아키텍처는 low-resolution feature maps과 이미지 크기에 따라 quadratic increase in complexity 때문에 dense vision tasks이나 입력 이미지 해상도의 범용 백본 네트워크로 사용하기에 적합하지 않음.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;38e035f8-2774-4c90-8c10-ba7caebf66ad&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;직접적인 upsampling 또는 deconvolution을 통한 object detection 과 semantic segmentation 의 dense vision tasks에 VIT 모델을 적용하는 몇 가지 연구가 있음[2, 78].&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;835b7b1f-1c7e-43d3-88d0-0f1d48d58962&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;더 나은 이미지 분류를 위해 ViT 아키텍처[69, 14, 27]를 수정하는 작업도 있음. Empirically, 이미지 분류에 관한 이러한 방법들 중에서 speed-accuracy trade-off을 달성하기 위해 Swin Transformer 아키텍처를 개발.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;1b986501-c44d-4a04-88c1-b945e58ee2df&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;비록 논문의 연구가 특별히 분류보다는 범용 성능에 초점을 맞추고 있음에도 불구하고, 또 다른 concurrent work[63]에서는 Transformers에서 multi-resolution feature maps을 구축하기 위한 유사한 사고 방식을 살펴봄. complexity는 여전히 이미지 크기에 quadratic 한 반면, 논문의 complexity는 linear이며 또한 locally 하게 작동하여 시각적 신호의 높은 상관 관계를 모델링하는 데 도움이 됨 [35, 24, 40].&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;5faa8697-08d7-4491-8f46-afd4dc4c7d9b&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;논문의 접근방식은 효율적이면서도 효과적이어서 COCO object detection와 ADE20K semantic segmentation 모두에서 state-of-the-art accuracy를 달성.&lt;/li&gt;
&lt;/ul&gt;
&lt;p id=&quot;09a82233-0a41-40db-a677-eea64dd48d3d&quot; class=&quot;&quot; data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;/details&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;c6f5917b-d17f-463a-97ad-5acb4f2c14a6&quot; class=&quot;&quot; data-ke-size=&quot;size26&quot;&gt;Method&lt;/h2&gt;
&lt;h3 id=&quot;03330731-70fb-469f-923c-7be4dbb5ac87&quot; class=&quot;&quot; data-ke-size=&quot;size23&quot;&gt;Overall Architecture&lt;/h3&gt;
&lt;figure id=&quot;9ab06d9c-181f-4827-8a35-e4ccf9b79b24&quot; class=&quot;image&quot;&gt;&lt;a href=&quot;https://i.imgur.com/IdxO6CZ.png[/img]&quot;&gt;&lt;img style=&quot;width: 1493px;&quot; src=&quot;https://i.imgur.com/IdxO6CZ.png[/img]&quot; /&gt;&lt;/a&gt;&lt;/figure&gt;
&lt;ul id=&quot;345a00ef-68af-49c7-8930-faf5d74a215d&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;그림 3에는 작은 버전(Swin-T)을 보여주는 Swin Transformer 아키텍처의 개요 소개.1) ViT와 같은 패치 분할 모듈에 의해 입력 RGB 이미지를 겹치지 않는 패치로 분할.
&lt;ul id=&quot;e9d7e27c-e8a7-4748-93c2-da9e2212ba20&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;각 패치는 &quot;token&quot;으로 처리되고 해당 feature는 raw pixel RGB 값의 연결로 설정됨. 구현 시, 4 &amp;times; 4의 패치 크기를 사용하므로 각 패치의 feature dimension는 4 &amp;times; 4 &amp;times; 3 = 48.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;86b82f26-7429-449f-88d4-d16743fc008f&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;이 raw-valued feature에 linear embedding layer가 적용되어 arbitrary dimension(C로 표시됨)로 투영됨.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;c559ebc7-aebb-4bf0-8386-b9c244a0c771&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Stage 1 : modified self-attention computation(Swin Transformer 블록)이 수정된 여러 트랜스포머 블록이 패치 토큰에 적용됨. Transformer 블록은 $&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;(H4&amp;times;W4)(\frac{H}{4} \times \frac{W}{4})&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1.217331em; vertical-align: -0.345em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mopen nulldelimiter&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mfrac&quot;&gt;&lt;span class=&quot;vlist-t vlist-t2&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.872331em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -2.6550000000000002em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;4&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.23em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;border-bottom-width: 0.04em;&quot; class=&quot;frac-line&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.394em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.08125em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;H&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-s&quot;&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.345em;&quot; class=&quot;vlist&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose nulldelimiter&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;&amp;times;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1.217331em; vertical-align: -0.345em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mopen nulldelimiter&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mfrac&quot;&gt;&lt;span class=&quot;vlist-t vlist-t2&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.872331em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -2.6550000000000002em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;4&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.23em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;border-bottom-width: 0.04em;&quot; class=&quot;frac-line&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.394em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.13889em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;W&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-s&quot;&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.345em;&quot; class=&quot;vlist&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose nulldelimiter&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;)$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;의 토큰을 유지하며 linear embedding과 함께 사용&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;69113d7a-6e74-49c4-b2af-97356279998f&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;계층적 표현을 생성하기 위해 네트워크의 깊이가 깊어질수록 patch merging layers에 의해 토큰 수가 감소.
&lt;ul id=&quot;9314070f-bab5-4416-a4b2-55031fb3c8fb&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;첫 번째 patch merging layer는 2 &amp;times; 2 인접한 patch 의 각 그룹의 특징을 연결하고 4C차원 concatenated features에 linear layer를 적용.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;af4e402c-5b25-427c-adf2-b94f863e5755&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;이렇게 하면 토큰 수가 2 &amp;times; 2 = 4의 배수(resolution 의 2&amp;times; 다운샘플링)로 감소하고 출력 dimension이 2C로 설정됨.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;323bdea9-0769-4be5-99d0-d6b44cfb424d&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;이후 $&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;H8&amp;times;W8\frac{H}{8} \times \frac{W}{8}&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1.217331em; vertical-align: -0.345em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mopen nulldelimiter&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mfrac&quot;&gt;&lt;span class=&quot;vlist-t vlist-t2&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.872331em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -2.6550000000000002em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;8&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.23em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;border-bottom-width: 0.04em;&quot; class=&quot;frac-line&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.394em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.08125em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;H&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-s&quot;&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.345em;&quot; class=&quot;vlist&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose nulldelimiter&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;&amp;times;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1.217331em; vertical-align: -0.345em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mopen nulldelimiter&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mfrac&quot;&gt;&lt;span class=&quot;vlist-t vlist-t2&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.872331em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -2.6550000000000002em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;8&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.23em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;border-bottom-width: 0.04em;&quot; class=&quot;frac-line&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.394em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.13889em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;W$ &lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;resolution을 유지하면서 feature transformation을 하기 위해 Swin Transformer 블록을 적용.
&lt;ul id=&quot;8eef671f-889a-4d1d-93b3-c36ea20148f0&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;이 patch merging 및 feature transformation의 첫 번째 블록을 &quot;Stage 2&quot;로 표시.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;3d892e17-1b82-4192-b8ec-564c3e172bf7&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;$$H16&amp;times;W16\frac{H}{16} \times \frac{W}{16}&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1.217331em; vertical-align: -0.345em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mopen nulldelimiter&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mfrac&quot;&gt;&lt;span class=&quot;vlist-t vlist-t2&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.872331em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -2.6550000000000002em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;16&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.23em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;border-bottom-width: 0.04em;&quot; class=&quot;frac-line&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.394em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.08125em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;H&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-s&quot;&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.345em;&quot; class=&quot;vlist&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose nulldelimiter&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;&amp;times;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1.217331em; vertical-align: -0.345em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mopen nulldelimiter&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mfrac&quot;&gt;&lt;span class=&quot;vlist-t vlist-t2&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.872331em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -2.6550000000000002em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;16&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.23em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;border-bottom-width: 0.04em;&quot; class=&quot;frac-line&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.394em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.13889em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;W$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-s&quot;&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.345em;&quot; class=&quot;vlist&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose nulldelimiter&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;과$&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;H32&amp;times;W32\frac{H}{32} \times \frac{W}{32}&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1.217331em; vertical-align: -0.345em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mopen nulldelimiter&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mfrac&quot;&gt;&lt;span class=&quot;vlist-t vlist-t2&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.872331em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -2.6550000000000002em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;32&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.23em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;border-bottom-width: 0.04em;&quot; class=&quot;frac-line&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.394em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.08125em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;H&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-s&quot;&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.345em;&quot; class=&quot;vlist&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose nulldelimiter&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;&amp;times;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1.217331em; vertical-align: -0.345em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mopen nulldelimiter&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mfrac&quot;&gt;&lt;span class=&quot;vlist-t vlist-t2&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.872331em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -2.6550000000000002em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;32&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.23em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;border-bottom-width: 0.04em;&quot; class=&quot;frac-line&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.394em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.13889em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;W&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-s&quot;&gt;​$$&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.345em;&quot; class=&quot;vlist&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose nulldelimiter&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;의 output resolutions 은 각각 'Stage 3'와 'Stage 4'로 두 차례 반복.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;c94c5e45-0346-465b-8516-f64ccf5e8d88&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;이러한 단계는 VGG [51] 및 ResNet [29]와 같은 일반적인 컨볼루션 네트워크의 것과 동일한 feature map resolutions으로 hierarchical representation을 공동으로 생성.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;3b958764-7188-4212-a148-988cbfcc50ee&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;그 결과, 제안된 아키텍처는 다양한 비전 작업을 위해 기존 방식에서 백본 네트워크를 편리하게 대체할 수 있어짐.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;5e19c9e2-cf23-4866-b9d6-a2742aad4b7a&quot; class=&quot;toggle&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;&lt;details&gt;
&lt;summary&gt;MLP layer&lt;/summary&gt;
&lt;p id=&quot;602c129c-8684-4df9-a85b-5f81cf544716&quot; class=&quot;&quot; data-ke-size=&quot;size16&quot;&gt;Fully connected - GELU - Fully connected&lt;/p&gt;
&lt;figure id=&quot;ded544e1-6ab9-4f3a-ba2e-a17a268252cd&quot; class=&quot;image&quot;&gt;&lt;a href=&quot;https://i.imgur.com/fyZrMeL.png[/img]&quot;&gt;&lt;img style=&quot;width: 902px;&quot; src=&quot;https://i.imgur.com/fyZrMeL.png[/img]&quot; /&gt;&lt;/a&gt;
&lt;figcaption&gt;reference : MLP Mixer&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/details&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;0d75519e-7780-4bca-803a-a5c56ede9ed9&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Swin Transformer block의 Swin Transformer는 Transformer block의 표준 multi-head self attention(MSA) 모듈을 Shift window를 기반으로 하는 모듈로 교체하고(섹션 3.2 참조) 다른 레이어를 동일하게 유지.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;4b943ade-4f92-49de-bf12-29a5a86d19b3&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;그림 3(b)에서 볼 수 있듯이, Swin Transformer 블록은 shifted window 기반 MSA 모듈로 구성되고 그 다음에 GELU 비선형성(수렴 속도 빠름) 을 사이에 둔 &lt;b&gt;이단 MLP&lt;/b&gt;로 구성됨.
&lt;figure id=&quot;d436ed4a-a70a-4cc0-99e8-8da9dd437daa&quot; class=&quot;image&quot;&gt;&lt;a href=&quot;https://i.imgur.com/VnZnwY1.png[/img]&quot;&gt;&lt;img style=&quot;width: 809px;&quot; src=&quot;https://i.imgur.com/VnZnwY1.png[/img]&quot; /&gt;&lt;/a&gt;&lt;/figure&gt;
&lt;ul id=&quot;f56ea754-dfd6-438a-a597-db2171c8d85e&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;각 MSA 모듈 및 각 MLP 앞에&lt;b&gt; LN(Layer Norm)&lt;/b&gt; 레이어가 적용되고 각 모듈 뒤에 residual connection이 적용됨.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;4dc40e70-b3c4-495e-b65a-3c4485a9e6c0&quot; data-ke-size=&quot;size23&quot;&gt;Shifted Window based Self-Attention&lt;/h3&gt;
&lt;ul id=&quot;35ab8cea-02b5-4eb6-a686-01dbb54e9bc1&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;표준 Transformer 아키텍처[61]와 이미지 분류에 대한 adaptation[19]은 토큰과 다른 모든 토큰 간의 관계가 계산되는 global self-attention를 수행함.
&lt;ul id=&quot;2eac6bff-24f0-4aff-a3a5-c4bff7005298&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;global computation은 토큰 수와 관련하여 &lt;b&gt;quadratic complexity&lt;/b&gt;으로 이어지며, dense prediction 또는 high-resolution 이미지를 나타내기 위해 엄청난 토큰 집합이 필요한 많은 비전 문제에 적합하지 않음.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;1c38b17b-3995-4835-99a4-2b2bc1d51ba4&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Self-attention in non-overlapped windows
&lt;ul id=&quot;7861fda7-9db1-4081-9625-0201dfbbfc12&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;효율적인 모델링을 위해 local windows의 self-attention를 계산할 것을 제안함.
&lt;ul id=&quot;8d8dd74d-2c60-4091-b2d6-1f38e69111ed&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;windows는 겹치지 않는 방식으로 이미지를 고르게 분할하도록 배열됨. 각 window에 M &amp;times; M 패치가 포함되어 있다고 가정할 때, global MSA 모듈의 계산 복잡도와 h &amp;times; w 패치의 이미지를 기반으로 하는 windows는
&lt;ul id=&quot;b77d10e2-4b33-44e4-afb6-47352eab2f51&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;$&amp;Omega;(MSA)=4hwC2+2(hw)2C\Omega(MSA) = 4hwC^2 + 2(hw)^2 C &lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1em; vertical-align: -0.25em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&amp;Omega;&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;(&lt;/span&gt;&lt;span style=&quot;margin-right: 0.05764em;&quot; class=&quot;mord mathnormal&quot;&gt;MS&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;A&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;)&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mrel&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.897438em; vertical-align: -0.08333em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;4&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;h&lt;/span&gt;&lt;span style=&quot;margin-right: 0.02691em;&quot; class=&quot;mord mathnormal&quot;&gt;w&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span style=&quot;margin-right: 0.07153em;&quot; class=&quot;mord mathnormal&quot;&gt;C&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.8141079999999999em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;2&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;+&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1.064108em; vertical-align: -0.25em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;h&lt;/span&gt;&lt;span style=&quot;margin-right: 0.02691em;&quot; class=&quot;mord mathnormal&quot;&gt;w&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;&lt;span class=&quot;mclose&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.8141079999999999em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;2&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.07153em;&quot; class=&quot;mord mathnormal&quot;&gt;C $&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;(1) global MSA&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;3c20784e-5724-456a-ba9a-2b3ad1ff749e&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;$&amp;Omega;(WMSA)=4hwC2+2M2hwC\Omega(WMSA) = 4hwC^2 + 2M^2hwC&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1em; vertical-align: -0.25em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&amp;Omega;&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;(&lt;/span&gt;&lt;span style=&quot;margin-right: 0.13889em;&quot; class=&quot;mord mathnormal&quot;&gt;W&lt;/span&gt;&lt;span style=&quot;margin-right: 0.05764em;&quot; class=&quot;mord mathnormal&quot;&gt;MS&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;A&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;)&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mrel&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.897438em; vertical-align: -0.08333em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;4&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;h&lt;/span&gt;&lt;span style=&quot;margin-right: 0.02691em;&quot; class=&quot;mord mathnormal&quot;&gt;w&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span style=&quot;margin-right: 0.07153em;&quot; class=&quot;mord mathnormal&quot;&gt;C&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.8141079999999999em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;2&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;+&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.8141079999999999em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal&quot;&gt;M&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.8141079999999999em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;2&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;h&lt;/span&gt;&lt;span style=&quot;margin-right: 0.07153em;&quot; class=&quot;mord mathnormal&quot;&gt;wC$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;(2) Shifted Window MSA
&lt;ul id=&quot;d7b499da-b53b-4a27-8339-3fd35495f237&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;여기서 (1)은 패치 번호&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;hwhw&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.69444em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;h&lt;/span&gt;&lt;span style=&quot;margin-right: 0.02691em;&quot; class=&quot;mord mathnormal&quot;&gt;w&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;에 quadratic이고 (2)는&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;MM&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.68333em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal&quot;&gt;M&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;이 고정되면 후자는 linear(기본적으로 &lt;b&gt;7&lt;/b&gt;로 설정됨).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;global self-attention computation은 일반적으로 large&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;hwhw&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.69444em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;h&lt;/span&gt;&lt;span style=&quot;margin-right: 0.02691em;&quot; class=&quot;mord mathnormal&quot;&gt;w&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;에 비해 unaffordable한 반면 window based self-attention은 확장 가능.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;8a3be13c-de4d-4a1a-a212-ef012094881d&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Shifted window partitioning in successive blocks
&lt;ul id=&quot;9c7f226d-2e86-4b18-a952-8b49b4f9507f&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;window-based self-attention module에는 window 간 연결이 부족하여 모델링 성능이 제한됨.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;02bd9a5e-8432-4982-901d-8878bc557fdc&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;non-overlapping windows의 효율적인 계산을 유지하면서 cross-window connections을 도입하기 위해 연속적인 SwinTransformer 블록의 두 분할 구성을 번갈아 사용하는 shifted window partitioning 방식 제안.&amp;nbsp;
&lt;ul id=&quot;5fe8b38f-82d6-451d-940c-ef0dedc405bf&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;그림 2에서 볼 수 있듯이, 첫 번째 모듈은 왼쪽 상단 픽셀에서 시작되는 regular window partitioning strategy을 사용하며, 8 &amp;times; 8 feature map은 크기가 4 &amp;times; 4 (M = 4)인 2 &amp;times; 2 windows로 균등하게 분할됨.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;91d9276f-9a3d-4800-8f9f-3d7f299b8d27&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;다음 모듈은 정기적으로 분할된 windows에서&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;([M2],[M2])([\frac{M}{2}], [\frac{M}{2}])&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1.217331em; vertical-align: -0.345em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;([&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mopen nulldelimiter&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mfrac&quot;&gt;&lt;span class=&quot;vlist-t vlist-t2&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.872331em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -2.6550000000000002em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;2&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.23em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;border-bottom-width: 0.04em;&quot; class=&quot;frac-line&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.394em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;M&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-s&quot;&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.345em;&quot; class=&quot;vlist&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose nulldelimiter&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;mpunct&quot;&gt;,&lt;/span&gt;&lt;span style=&quot;margin-right: 0.16666666666666666em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mopen nulldelimiter&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mfrac&quot;&gt;&lt;span class=&quot;vlist-t vlist-t2&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.872331em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -2.6550000000000002em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;2&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.23em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;border-bottom-width: 0.04em;&quot; class=&quot;frac-line&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.394em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;M&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-s&quot;&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.345em;&quot; class=&quot;vlist&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose nulldelimiter&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;])&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;픽셀로 windows를 대체하여 이전 계층에서 shifted window 구성을 채택.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;13710c4c-20e8-4228-af95-843025ba4a35&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;shifted window partitioning approach를 사용하면 연속적인 Swin Transformer 블록이 다음과 같이 계산됨.
&lt;ul id=&quot;56741bcd-b0b3-47f0-86b3-f5494801b240&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;$z^=W&amp;minus;MSA(LN(zl&amp;minus;1))+zl&amp;minus;1\hat{z} = W-MSA(LN(z^{l-1}))+z^{l-1}&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.69444em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord accent&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.69444em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.04398em;&quot; class=&quot;mord mathnormal&quot;&gt;z&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;left: -0.19444em;&quot; class=&quot;accent-body&quot;&gt;&lt;span class=&quot;mord&quot;&gt;^&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mrel&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.76666em; vertical-align: -0.08333em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.13889em;&quot; class=&quot;mord mathnormal&quot;&gt;W&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;&amp;minus;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1.0991079999999998em; vertical-align: -0.25em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.05764em;&quot; class=&quot;mord mathnormal&quot;&gt;MS&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;A&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;L&lt;/span&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal&quot;&gt;N&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span style=&quot;margin-right: 0.04398em;&quot; class=&quot;mord mathnormal&quot;&gt;z&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.8491079999999999em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.01968em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;l&lt;/span&gt;&lt;span class=&quot;mbin mtight&quot;&gt;&amp;minus;&lt;/span&gt;&lt;span class=&quot;mord mtight&quot;&gt;1&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;))&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;+&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.8491079999999999em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span style=&quot;margin-right: 0.04398em;&quot; class=&quot;mord mathnormal&quot;&gt;z&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.8491079999999999em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.01968em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;l&lt;/span&gt;&lt;span class=&quot;mbin mtight&quot;&gt;&amp;minus;&lt;/span&gt;&lt;span class=&quot;mord mtight&quot;&gt;1$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;cdab1e82-8bd9-4f77-8a71-94e6654f370a&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;$zl=MLP(LN(z^l))+z^lz^l = MLP(LN(\hat{z}^l)) + \hat{z}^l&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.849108em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span style=&quot;margin-right: 0.04398em;&quot; class=&quot;mord mathnormal&quot;&gt;z&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.849108em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.01968em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;l&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mrel&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1.099108em; vertical-align: -0.25em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal&quot;&gt;M&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;L&lt;/span&gt;&lt;span style=&quot;margin-right: 0.13889em;&quot; class=&quot;mord mathnormal&quot;&gt;P&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;L&lt;/span&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal&quot;&gt;N&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mord accent&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.69444em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.04398em;&quot; class=&quot;mord mathnormal&quot;&gt;z&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;left: -0.19444em;&quot; class=&quot;accent-body&quot;&gt;&lt;span class=&quot;mord&quot;&gt;^&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.849108em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.01968em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;l&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;))&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;+&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.849108em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mord accent&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.69444em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.04398em;&quot; class=&quot;mord mathnormal&quot;&gt;z&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;left: -0.19444em;&quot; class=&quot;accent-body&quot;&gt;&lt;span class=&quot;mord&quot;&gt;^&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.849108em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.01968em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;l $&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;,&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;138d3cbc-3f3c-4981-bd55-7674cfdc0952&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;$z^l+1=SW&amp;minus;MSA(LN(zl))+zl\hat{z}^{l+1} = SW-MSA(LN(z^l))+z^l&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.8491079999999999em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mord accent&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.69444em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.04398em;&quot; class=&quot;mord mathnormal&quot;&gt;z&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;left: -0.19444em;&quot; class=&quot;accent-body&quot;&gt;&lt;span class=&quot;mord&quot;&gt;^&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.8491079999999999em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.01968em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;l&lt;/span&gt;&lt;span class=&quot;mbin mtight&quot;&gt;+&lt;/span&gt;&lt;span class=&quot;mord mtight&quot;&gt;1&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mrel&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.76666em; vertical-align: -0.08333em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.05764em;&quot; class=&quot;mord mathnormal&quot;&gt;S&lt;/span&gt;&lt;span style=&quot;margin-right: 0.13889em;&quot; class=&quot;mord mathnormal&quot;&gt;W&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;&amp;minus;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1.099108em; vertical-align: -0.25em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.05764em;&quot; class=&quot;mord mathnormal&quot;&gt;MS&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;A&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;L&lt;/span&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal&quot;&gt;N&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span style=&quot;margin-right: 0.04398em;&quot; class=&quot;mord mathnormal&quot;&gt;z&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.849108em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.01968em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;l&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;))&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;+&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.849108em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span style=&quot;margin-right: 0.04398em;&quot; class=&quot;mord mathnormal&quot;&gt;z&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.849108em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.01968em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;l $&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;,&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;9bec648e-bd73-4922-a0b9-4645b96c88e2&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;$zl+1=MLP(LN(z^l+1))+z^l+1z^{l+1} = MLP(LN(\hat{z}^{l+1})) + \hat{z}^{l+1}&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.8491079999999999em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span style=&quot;margin-right: 0.04398em;&quot; class=&quot;mord mathnormal&quot;&gt;z&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.8491079999999999em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.01968em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;l&lt;/span&gt;&lt;span class=&quot;mbin mtight&quot;&gt;+&lt;/span&gt;&lt;span class=&quot;mord mtight&quot;&gt;1&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mrel&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1.0991079999999998em; vertical-align: -0.25em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal&quot;&gt;M&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;L&lt;/span&gt;&lt;span style=&quot;margin-right: 0.13889em;&quot; class=&quot;mord mathnormal&quot;&gt;P&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;L&lt;/span&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal&quot;&gt;N&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mord accent&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.69444em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.04398em;&quot; class=&quot;mord mathnormal&quot;&gt;z&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;left: -0.19444em;&quot; class=&quot;accent-body&quot;&gt;&lt;span class=&quot;mord&quot;&gt;^&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.8491079999999999em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.01968em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;l&lt;/span&gt;&lt;span class=&quot;mbin mtight&quot;&gt;+&lt;/span&gt;&lt;span class=&quot;mord mtight&quot;&gt;1&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;))&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;+&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.8491079999999999em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mord accent&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.69444em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.04398em;&quot; class=&quot;mord mathnormal&quot;&gt;z&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;left: -0.19444em;&quot; class=&quot;accent-body&quot;&gt;&lt;span class=&quot;mord&quot;&gt;^&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.8491079999999999em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.01968em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;l&lt;/span&gt;&lt;span class=&quot;mbin mtight&quot;&gt;+&lt;/span&gt;&lt;span class=&quot;mord mtight&quot;&gt;1$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt; (3)
&lt;ul id=&quot;353db490-a342-4e64-b560-bcc05a3b207b&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;여기서 $&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;z^\hat{z}&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.69444em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord accent&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.69444em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.04398em;&quot; class=&quot;mord mathnormal&quot;&gt;z&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;left: -0.19444em;&quot; class=&quot;accent-body&quot;&gt;&lt;span class=&quot;mord&quot;&gt;^&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;과&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;zlz^l&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.849108em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span style=&quot;margin-right: 0.04398em;&quot; class=&quot;mord mathnormal&quot;&gt;z&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.849108em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.01968em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;l$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt; 은 각각 블록&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;ll&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.69444em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.01968em;&quot; class=&quot;mord mathnormal&quot;&gt;l&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;에 대한 (S)W-MSA 모듈과 MLP 모듈의 출력 기능을 나타냅니다. W-MSA 및 SW-MSA는 각각 일반 및 shifted window partitioning 구성을 사용한 window based multi-head self-attention를 나타냄.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li id=&quot;64faf894-6027-4636-b7b5-62cbdbcb70fe&quot; class=&quot;image&quot;&gt;&lt;a href=&quot;https://i.imgur.com/bLvZYuc.png[/img]&quot;&gt;&lt;img style=&quot;width: 725px;&quot; src=&quot;https://i.imgur.com/bLvZYuc.png[/img]&quot; /&gt;&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;3ffe8fb8-b8aa-4a54-aa72-64bfc85549a0&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;shifted window partitioning 접근방식은 이전 계층에서 인접한 non-overlapping windows 사이의 connections을 도입하며, 표 4와 같이 image classification, object detection, and semantic segmentation,에 효과적인 것으로 확인됨.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;fe1432af-f22a-4903-ae6a-20ad8a5640d5&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Efficient batch computation for shifted configuration
&lt;ul id=&quot;3a3af16f-10a1-4a31-9930-f4518986cedd&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;shifted window partitioning의 문제는 이동 구성의 $&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;[hM]&amp;times;[wM]&amp;nbsp;&amp;nbsp;to&amp;nbsp;([hM]+1)&amp;times;([wM]+1)[\frac{h}{M}] \times [\frac{w}{M}] ~~{\rm to}~ ([\frac{h}{M}] + 1) \times ([\frac{w}{M}]+1) &lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1.2251079999999999em; vertical-align: -0.345em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mopen nulldelimiter&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mfrac&quot;&gt;&lt;span class=&quot;vlist-t vlist-t2&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.8801079999999999em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -2.6550000000000002em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;M&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.23em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;border-bottom-width: 0.04em;&quot; class=&quot;frac-line&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.394em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span class=&quot;mord mathnormal mtight&quot;&gt;h&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-s&quot;&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.345em;&quot; class=&quot;vlist&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose nulldelimiter&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;]&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;&amp;times;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1.2251079999999999em; vertical-align: -0.345em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mopen nulldelimiter&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mfrac&quot;&gt;&lt;span class=&quot;vlist-t vlist-t2&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.695392em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -2.6550000000000002em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;M&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.23em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;border-bottom-width: 0.04em;&quot; class=&quot;frac-line&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.394em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.02691em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;w&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-s&quot;&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.345em;&quot; class=&quot;vlist&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose nulldelimiter&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;mspace nobreak&quot;&gt;&amp;nbsp;&lt;/span&gt;&lt;span class=&quot;mspace nobreak&quot;&gt;&amp;nbsp;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mord mathrm&quot;&gt;to&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mspace nobreak&quot;&gt;&amp;nbsp;&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;([&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mopen nulldelimiter&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mfrac&quot;&gt;&lt;span class=&quot;vlist-t vlist-t2&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.8801079999999999em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -2.6550000000000002em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;M&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.23em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;border-bottom-width: 0.04em;&quot; class=&quot;frac-line&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.394em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span class=&quot;mord mathnormal mtight&quot;&gt;h&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-s&quot;&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.345em;&quot; class=&quot;vlist&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose nulldelimiter&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;]&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;+&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1em; vertical-align: -0.25em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;)&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;&amp;times;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1.095em; vertical-align: -0.345em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;([&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mopen nulldelimiter&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mfrac&quot;&gt;&lt;span class=&quot;vlist-t vlist-t2&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.695392em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -2.6550000000000002em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;M&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.23em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;border-bottom-width: 0.04em;&quot; class=&quot;frac-line&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.394em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.02691em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;w&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-s&quot;&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.345em;&quot; class=&quot;vlist&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose nulldelimiter&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;]&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;+&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1em; vertical-align: -0.25em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;)$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;에서 더 많은 window 가 발생하며 일부 window는 $&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;M&amp;times;M$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;보다 작음.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;83cc55b0-558d-41a3-8e0a-7af8a3fb1b39&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;naive solution은 size of $&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;M&amp;times;M$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&amp;nbsp;보다 작은windows는 pad 하고 a attention을 계산할 때 padded values을 mask out. regular partitioning의 the number of windows가 작을 때,
&lt;ul id=&quot;a7bdbda6-7a9c-4cdd-9102-680f704a5227&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;예) 2 &amp;times; 2일 때, 이 naive solution을 사용한 증가된 계산은 상당함. (2 &amp;times; 2 &amp;rarr; 3 &amp;times; 3 , 이는 9/4 = 2.25 배 더 큼.)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;e042e408-535b-46d7-8e8a-1de055fd41ba&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;여기서, 논문은 그림 4와 같이 top-left direction으로 cyclic-shifting함으로써 보다 more efficient batch computation approach를 제안.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;3d044070-acc7-4fef-9875-85e02ac9ebc4&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;이동 후에는 batched window가 feature map에 인접하지 않은 여러 sub-window(A : 색 4개 조합)로 구성될 수 있으므로 masking mechanism을 사용하여 각 하위 window 내에서 self-attention computation을 제한.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;eca77f01-d019-4741-8c02-f8c1951fbd8d&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;cyclic-shift을 사용하면 batched windows의 수가 regular window partitioning의 window 수와 동일하게 유지되므로 효율적이기도 함. 이 접근법의 low latency는 표 5에 나와 있음.&lt;/li&gt;
&lt;/ul&gt;
&lt;figure id=&quot;910c0f7d-db12-48a3-9bc7-4901c765520e&quot; class=&quot;image&quot;&gt;&lt;a href=&quot;https://i.imgur.com/k8O8X96.png[/img]&quot;&gt;&lt;img style=&quot;width: 741px;&quot; src=&quot;https://i.imgur.com/k8O8X96.png[/img]&quot; /&gt;&lt;/a&gt;
&lt;figcaption&gt;그림 4. shifted window partitioning에서 self-attention를 기울일 수 있는 효율적인 batch computation approach을 보여줍니다.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;057610f3-d13b-4b2d-ae85-92712c823fae&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Relative position bias
&lt;ul id=&quot;bbe9b49f-19a1-4476-acc3-172166286706&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;self-attention computation에서는 computing similarity 에 relative position bias $&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;B&amp;isin;RM2&amp;times;M2B \in \mathbb{R}^{M^2 &amp;times; M^2}&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.72243em; vertical-align: -0.0391em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.05017em;&quot; class=&quot;mord mathnormal&quot;&gt;B&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mrel&quot;&gt;&amp;isin;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.9869199999999999em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mord mathbb&quot;&gt;R&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.9869199999999999em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;M&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.8913142857142857em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -2.931em; margin-right: 0.07142857142857144em;&quot;&gt;&lt;span style=&quot;height: 2.5em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size3 size1 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;2&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mbin mtight&quot;&gt;&amp;times;&lt;/span&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;M&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.8913142857142857em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -2.931em; margin-right: 0.07142857142857144em;&quot;&gt;&lt;span style=&quot;height: 2.5em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size3 size1 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;2$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;를 포함시켜 [48, 1, 31, 32]$&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;Attention(Q,K,V)=SoftMax(QKT/d+B)VAttention(Q, K, V) = SoftMax(QK^T / \sqrt{d} + B) V&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1em; vertical-align: -0.25em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;A&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;tt&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;e&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;n&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;t&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;i&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;o&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;n&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;Q&lt;/span&gt;&lt;span class=&quot;mpunct&quot;&gt;,&lt;/span&gt;&lt;span style=&quot;margin-right: 0.16666666666666666em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.07153em;&quot; class=&quot;mord mathnormal&quot;&gt;K&lt;/span&gt;&lt;span class=&quot;mpunct&quot;&gt;,&lt;/span&gt;&lt;span style=&quot;margin-right: 0.16666666666666666em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.22222em;&quot; class=&quot;mord mathnormal&quot;&gt;V&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;)&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mrel&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1.18222em; vertical-align: -0.25em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.05764em;&quot; class=&quot;mord mathnormal&quot;&gt;S&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;o&lt;/span&gt;&lt;span style=&quot;margin-right: 0.10764em;&quot; class=&quot;mord mathnormal&quot;&gt;f&lt;/span&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal&quot;&gt;tM&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;a&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;Q&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span style=&quot;margin-right: 0.07153em;&quot; class=&quot;mord mathnormal&quot;&gt;K&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.8413309999999999em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.13889em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;T&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;/&lt;/span&gt;&lt;span class=&quot;mord sqrt&quot;&gt;&lt;span class=&quot;vlist-t vlist-t2&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.93222em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3em;&quot; class=&quot;svg-align&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;padding-left: 0.833em;&quot; class=&quot;mord&quot;&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;d&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -2.89222em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;min-width: 0.853em; height: 1.08em;&quot; class=&quot;hide-tail&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-s&quot;&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.10777999999999999em;&quot; class=&quot;vlist&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;+&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1em; vertical-align: -0.25em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.05017em;&quot; class=&quot;mord mathnormal&quot;&gt;B&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;)&lt;/span&gt;&lt;span style=&quot;margin-right: 0.22222em;&quot; class=&quot;mord mathnormal&quot;&gt;V$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;를 따름.
&lt;ul id=&quot;21439eea-f1f2-4f26-b70b-803091b67c17&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;여기서 $&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;Q,K,V&amp;isin;RM2&amp;times;dQ, K, V \in \mathbb{R}^{M^2&amp;times;d}&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.8777699999999999em; vertical-align: -0.19444em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;Q&lt;/span&gt;&lt;span class=&quot;mpunct&quot;&gt;,&lt;/span&gt;&lt;span style=&quot;margin-right: 0.16666666666666666em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.07153em;&quot; class=&quot;mord mathnormal&quot;&gt;K&lt;/span&gt;&lt;span class=&quot;mpunct&quot;&gt;,&lt;/span&gt;&lt;span style=&quot;margin-right: 0.16666666666666666em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.22222em;&quot; class=&quot;mord mathnormal&quot;&gt;V&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mrel&quot;&gt;&amp;isin;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.9869199999999999em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mord mathbb&quot;&gt;R&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.9869199999999999em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;M&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.8913142857142857em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -2.931em; margin-right: 0.07142857142857144em;&quot;&gt;&lt;span style=&quot;height: 2.5em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size3 size1 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;2&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;mbin mtight&quot;&gt;&amp;times;&lt;/span&gt;&lt;span class=&quot;mord mathnormal mtight&quot;&gt;d$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;는 쿼리, 키 및 값 매트릭스,&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;d&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;는 query/key dimension, $&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;M^2$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;는 창에 있는 패치 수. 각 축을 따라 상대적인 위치가 $&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;[&amp;minus;M+1,M&amp;minus;1][-M + 1, M -1] &lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1em; vertical-align: -0.25em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mopen&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&amp;minus;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal&quot;&gt;M&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;+&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.8777699999999999em; vertical-align: -0.19444em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;mpunct&quot;&gt;,&lt;/span&gt;&lt;span style=&quot;margin-right: 0.16666666666666666em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal&quot;&gt;M&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;&amp;minus;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 1em; vertical-align: -0.25em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;mclose&quot;&gt;]$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;범위에 있으므로 작은 크기의 바이어스 행렬$&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;B^&amp;isin;R(2M&amp;minus;1)&amp;times;(2M&amp;minus;1)\hat{B} \in \mathbb{R}^{(2M - 1)&amp;times;(2M - 1)}&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.9858699999999999em; vertical-align: -0.0391em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord accent&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.9467699999999999em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.05017em;&quot; class=&quot;mord mathnormal&quot;&gt;B&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;top: -3.25233em;&quot;&gt;&lt;span style=&quot;height: 3em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span style=&quot;left: -0.16666em;&quot; class=&quot;accent-body&quot;&gt;&lt;span class=&quot;mord&quot;&gt;^&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mrel&quot;&gt;&amp;isin;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2777777777777778em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.8879999999999999em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&lt;span class=&quot;mord mathbb&quot;&gt;R&lt;/span&gt;&lt;span class=&quot;msupsub&quot;&gt;&lt;span class=&quot;vlist-t&quot;&gt;&lt;span class=&quot;vlist-r&quot;&gt;&lt;span style=&quot;height: 0.8879999999999999em;&quot; class=&quot;vlist&quot;&gt;&lt;span style=&quot;top: -3.063em; margin-right: 0.05em;&quot;&gt;&lt;span style=&quot;height: 2.7em;&quot; class=&quot;pstrut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;sizing reset-size6 size3 mtight&quot;&gt;&lt;span class=&quot;mord mtight&quot;&gt;&lt;span class=&quot;mopen mtight&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mord mtight&quot;&gt;2&lt;/span&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;M&lt;/span&gt;&lt;span class=&quot;mbin mtight&quot;&gt;&amp;minus;&lt;/span&gt;&lt;span class=&quot;mord mtight&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;mclose mtight&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;mbin mtight&quot;&gt;&amp;times;&lt;/span&gt;&lt;span class=&quot;mopen mtight&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mord mtight&quot;&gt;2&lt;/span&gt;&lt;span style=&quot;margin-right: 0.10903em;&quot; class=&quot;mord mathnormal mtight&quot;&gt;M&lt;/span&gt;&lt;span class=&quot;mbin mtight&quot;&gt;&amp;minus;&lt;/span&gt;&lt;span class=&quot;mord mtight&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;mclose mtight&quot;&gt;)$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;의 값을&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;B&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;에서 취한다.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;01802cc1-73e2-4600-9dd1-6526429b81fd&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;표 4에서와 같이 이러한 bias 항이 없거나 absolute position embedding을 사용하는 논문에 비해 현저한 개선을 관찰.
&lt;ul id=&quot;05bd7957-9797-4a5b-9592-71fcdf60cc1e&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;[19]에서와 같이 absolute position embedding을 입력에 추가하면 성능이 약간 저하되므로 구현에서 채택되지 않음.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;17feb923-86fd-4a29-98a7-7dc980ace14b&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;pre-training에서 학습된 relative position bias을 사용하여 bi-cubic interpolation을 통해 다른 window size로 fine-tuning을 위한 모델을 초기화할 수도 있음[19, 60].&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;e1772f5b-71dd-47de-85f1-322177ce3021&quot; data-ke-size=&quot;size23&quot;&gt;Architecture Variants&lt;/h3&gt;
&lt;ul id=&quot;1742de20-9b2a-4b1d-88cc-fc4dab52c8d2&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;ViT-B/DeiT-B와 유사한 모델 크기와 계산 복잡성을 갖도록 Swin-B라는 기본 모델을 구축.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;6cbab228-0da9-4c74-a64e-2d9f77ec33d1&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;또한 모델 크기 약 0.25배, 0.5배, 2배인 Swin-T, Swin-S, Swin-L을 소개.
&lt;ul id=&quot;1e5a5070-0b42-44ca-a07e-7ae21f4c33a2&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;Swin-T와 Swin-S의 복잡도는 각각 ResNet-50(DeiT-S)과 ResNet-101과 유사.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;e2ae1083-dd7a-4b4d-8313-b0dc2b816549&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;window size는 기본적으로 M = 7로 설정.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;3a8099d5-6291-48d9-80e5-a89a1e16aa7a&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;모든 실험에서 각 head의 query dimension은 d = 32이고, 각 MLP의 확장 레이어는 &amp;alpha; = 4.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;06d9bc5c-cbf0-4506-aa90-b782dfb4579c&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;이러한 모델 모델의 아키텍처 하이퍼 파라미터는 다음과 같습니다.
&lt;ul id=&quot;692388b3-659f-4ba8-9ecf-d446a5c0d98a&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;Swin-T: C = 96, layer numbers = {2, 2, 6, 2}&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;0aa00669-0a06-4e87-8a62-8ff7bf979911&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;Swin-S: C = 96, layer numbers ={2, 2, 18, 2}&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;51dcde25-81e8-4a3a-83c4-2d7a9e82cfbd&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;Swin-B: C = 128, layer numbers ={2, 2, 18, 2}&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;44105d68-34ba-4e01-9a18-5316a7b25ef3&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;Swin-L: C = 192, layer numbers ={2, 2, 18, 2}&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;95b3ee24-648d-46dc-9257-e8ad1c40f3f3&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;여기서 C는 첫 번째 단계에서 숨겨진 레이어의 채널 number.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;9392f878-04ef-43f1-9320-c0ef4d45f3cd&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;ImageNet 이미지 분류에 대한 모델 크기, 이론적 계산 복잡성(FLOP) 및 모델 변형 처리량은 표 1.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;ed3c285a-ab03-4d04-aeae-d18a5420932c&quot; data-ke-size=&quot;size26&quot;&gt;Experiments&lt;/h2&gt;
&lt;ul id=&quot;bcf50b92-804d-42c4-a94c-bedce8fd7404&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;ImageNet-1K image classification[18], COCO object detection[42] 및 ADE20K semantic segmentation[80]에 대한 실험을 수행.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;4997c763-216d-4f7c-9af3-8d56a104ac11&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;다음에서는 먼저 제안된 Swin Transformer 아키텍처를 세 가지 작업에 대한 이전의 SOTA와 비교. 그런 다음 Swin Transformer의 중요한 디자인 요소를 완화.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;72b0b825-3faf-4f85-b02d-a62a71020ba3&quot; data-ke-size=&quot;size23&quot;&gt;&amp;nbsp;Image Classiﬁcation on ImageNet-1K&lt;/h3&gt;
&lt;ul id=&quot;09fee897-9748-4782-8add-16bfb3fbe19b&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Settings
&lt;ul id=&quot;a49bbce8-58eb-456d-9ec9-f18b7001832f&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;이미지 분류를 위해 1,000개의 클래스에서 128M개의 training 이미지와 50K valid 이미지가 포함된 ImageNet-1K[18]에서 제안된 Swin Transformer를 벤치마크함.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;4c48322d-4edb-4810-b982-b6cfd9b93b27&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;single crop에서 top-1 accuracy가 보고.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;18f7c343-9ded-489f-a131-759a017de2f2&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;다음 두 가지 교육 설정을 고려합니다.
&lt;ul id=&quot;ea4f1ad1-05b5-42b3-a65a-80383bd5d123&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;Regular ImageNet-1K training.
&lt;ul id=&quot;36ad0884-5380-45d1-89a3-6f93cfaedf38&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;이 설정은 대부분 [60]을 따릅니다. cosine decay learning ratescheduler와 20epoch의 linear warm-up을 사용하여 300epoch에 대해 AdamW[36] optimizer를 사용.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;94700ec1-c5b1-42d6-977e-37c7383ab3e4&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;배치 크기 1024, 초기 학습 속도 0.001 및 weight decay 0.05가 사용됨.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;41ac2803-beb4-4110-8939-17257b0dcf25&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;[30] 및 [44]의 성능을 향상시키지 않는 반복적인 augmentation을 제외하고 [60]의 대부분의 augmentation and regularization 전략을 training에 포함시킴. 이는 ViT의 training을 안정화하는 데 있어 반복적인 augmentation이 중요한 [60]과 반대되는 점에 유의.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;1258aa64-06e0-4a8e-a45f-cdf99425dabe&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;ImageNet-22K에 대한 Pre-training 및 ImageNet-1K에 대한 fine-tuning 또한 14.2 million 개의 이미지와 22K 클래스를 포함하는 대규모 ImageNet-22K 데이터셋에 대한 pre-train도 실시.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;d65b0002-e1b9-4a76-9679-e74b69daa494&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;5-epoch linear warm-up0과 함께 linear decay learning rate scheduler를 사용하여 60 epochs에 AdamW optimizer를 사용.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;1c4559be-f113-4a40-99d8-1ded81d816fe&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;배치 크기 4096, 초기 학습률 0.001 및 weight decay 0.01이 사용. ImageNet-1K fine-tuning에서는 배치 크기 1024, 일정한 학습 속도 10^-5, weight decay 10^-8로 30 epochs 모델을 training.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;25c3463d-dbf7-468c-9150-a6768eef4364&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Results with regular ImageNet-1K training
&lt;ul id=&quot;f8eda4a9-d482-43b5-a52f-5a3671d6d264&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;표 1(a)은 regular ImageNet-1K training을 사용하여 Transformer-based 와 ConvNet-based 모두를 포함한 다른 백본과 비교state-of-the-art Transformerbased.&lt;/li&gt;
&lt;/ul&gt;
&lt;figure id=&quot;36928311-a139-4c27-a9c1-8ac692d14632&quot; class=&quot;image&quot;&gt;&lt;a href=&quot;https://i.imgur.com/KDbudTK.png[/img]&quot;&gt;&lt;img style=&quot;width: 744px;&quot; src=&quot;https://i.imgur.com/KDbudTK.png[/img]&quot; /&gt;&lt;/a&gt;
&lt;figcaption&gt;table1. Comparison of different backbones on ImageNet-1K classification. Throughput is measured using the GitHub repository of [65] and a V100 GPU, following [60].&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;ul id=&quot;9b9fd4bb-fdd4-41d3-94b4-4866e848fa2f&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;이전의 state-of-the-art Transformerbased architecture(예: DeiT[60])와 비교했을 때, Swin Transformers는$&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;224^2$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;를 사용한 Dei-S(79.8%)에 비해 복잡성이 비슷한 DeiT 아키텍처를 눈에 띄게 능가.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;330b14c8-210e-477f-9a61-af8f79c692f5&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;최신 ConvNet, 즉 Reg-Net[47] 및 EfficientNet[57]과 비교했을 때, Swin Transformer는 speed-accuracy trade-off를 약간 더 잘 달성.
&lt;ul id=&quot;1ad81e28-cc61-4e06-a42d-6168e05a6cfb&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;RegNet [47]과 EfficientNet [57]은 철저한 아키텍처 search를 통해 확보되지만 제안된 Swin Transformer는 표준 트랜스포머에서 채택되어 추가 개선 가능성이 크다는 점에 주목.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;b8b84a2f-b058-49da-be56-1d00e52dc457&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Results with ImageNet-22K pre-training
&lt;ul id=&quot;cfafff32-23bd-4270-a7bf-05b4d8a94a96&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;ImageNet-22K에서 대용량 Swin-B 및 Swin-L도 pretrain.
&lt;ul id=&quot;d45eff53-7340-4d2f-b77a-3839ba9eb7a0&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;ImageNet-1K 영상 분류에서 fine-tuned 결과는 표 1(b).&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;82acb981-5c6f-4057-b82a-5eb7e1c39e2b&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;Swin-B의 경우 ImageNet-22K pretrain은 처음부터 ImageNet-1K에 대한 train에 비해 1.8%~1.9% 향상.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;6c58126c-a21f-4383-9679-1f48e2015877&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;ImageNet-22K pre-training에 대한 이전의 최상 결과와 비교했을 때, 논문의 모델은speed-accuracy trade-offs 측면에서 훨씬 더 나음.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;741d758d-c652-4095-9a4c-a9a9c52aebca&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;Swin-B는 86.0%의 top-1 accuracy를 얻어 inference throughput(84.7 vs. 85.9 영상/초)이 비슷한 ViT보다 2.0% 높고 FLOP(47.0G vs. 55.4G)가 약간 낮다.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;551591e2-93b4-4635-8e97-c63e55d8f127&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;대형 Swin-L 모델은 86.4%의 Top-1 정확도를 달성하여 Swin-B 모델보다 약간 우수합니다.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;1a48c16c-a724-4399-ab8e-1655913df032&quot; data-ke-size=&quot;size23&quot;&gt;Object Detection on COCO&lt;/h3&gt;
&lt;ul id=&quot;8fff0b44-a861-4299-b31b-7e9bce038bae&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Settings
&lt;ul id=&quot;8cf71ca0-a534-43a4-aa65-298b538114fb&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;Object detection and instance segmentation은 118K training, 5K validation 및 20K test-dev images가 포함된 COCO 2017에서 수행.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;1c222f39-7370-41d0-9e47-01db038382a3&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;validation 세트를 사용하여 ablation study가 수행되며, test-dev 시 system-level comparison가 보고.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;d68c6d9d-c9d9-485c-ae2c-2cbfb71f3164&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;ablation study를 위해, 우리는 네 가지 일반적인 object detection frameworks: Cascade Mask R-CNN [28, 5], ATSS [76], RepPoints v2 [11], and Sparse RCNN [55] in mmdetection[9].를 고려.
&lt;ul id=&quot;cddd9d86-91b3-4c75-a513-dde1d3e5d000&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;이 네 가지 프레임워크에 대해, 우리는 same settings: multi-scale training[7, 55](짧은 쪽이 480에서 800 사이인 반면 긴 쪽이 최대 1333 사이인 입력의 크기 조정), Adam W[43] optimizer(초기 학습 속도 0.0001, weight decay 0.05, 배치 크기 16), 3x 스케줄(36 epoches)을 활용.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;0f60f60d-2fd6-4df4-b3f1-6c4a5b7683c9&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;system-level comparison를 위해, 우리는 improved HTC[8] (HTC++로 표시됨), instaboost[21], 보다 강력한 multi-scale training [6], 6x schedule(72 epoch), soft-NMS[4] 및 ImageNet-22K pre-trained model을 초기화로 채택.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;c68c1908-b43c-4758-aeee-6f409c48d1c8&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;우리는 Swin Transformer를 표준 Con-vNets(예: ResNe(X)t) 및 이전 Transformer 네트워크(예: DeiT)와 비교.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;1885d11e-a2d0-4836-8f7f-c20116d4e8c6&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;비교는 다른 설정이 변경되지 않은 백본만 변경하여 수행.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;fd53e759-2eec-4ac7-802c-b006810bfd07&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;Swin Transformer 및 ResNe(X)t는 hierarchical feature maps으로 인해 위의 모든 프레임워크에 직접 적용할 수 있지만 DeiT는 피쳐 맵의 단일 해상도만 생성하며 직접 적용할 수 없음.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;a19abacf-8060-4ec7-8023-68bbb279fd1e&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;공정한 비교를 위해, 우리는 deconvolution 레이어를 사용하여 DeiT에 대한 hierarchical feature maps을 구성하기 위해 [78]을 따릅니다.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;f29fbf13-ea15-461b-80b1-214d99c815c9&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Comparison to ResNe(X)t&amp;nbsp;
&lt;ul id=&quot;80813758-5545-4da4-b052-7694b74b2a02&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;표 2(a)는 네 개의 object detection frameworks에 대한 Swin-T 및 ResNet-50의 결과를 나열.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;f7f213b5-6cb3-40de-9a0a-d14f05f418e9&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;Swin-T 아키텍처는 ResNet-50에 비해 일관된 +3.4~4.2box AP 이점을 제공하며, 모델 크기, FLOP 및 대기 시간이 약간 더 큼.&lt;/li&gt;
&lt;/ul&gt;
&lt;figure id=&quot;880e8fc4-d343-463a-8ee4-c7eddf6aaaa1&quot; class=&quot;image&quot;&gt;&lt;a href=&quot;https://i.imgur.com/sBq7Eff.png[/img]&quot;&gt;&lt;img style=&quot;width: 481px;&quot; src=&quot;https://i.imgur.com/sBq7Eff.png[/img]&quot; /&gt;&lt;/a&gt;&lt;/figure&gt;
&lt;ul id=&quot;b9a6713c-69a2-4f7a-94e5-bf0161470078&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;표 2(b)는 Cascade Mask R-CNN을 사용하여 서로 다른 모델 용량에서 Swin Transformer와 ResNe(X)를 비교.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;864febb2-0085-437c-acc4-bd19a113ab81&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;Swin Transformer는 ResNext에 비해 +3.6 box AP 및 +3.3 mask AP의 높은 detection accuracy를 달성.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;77eec721-8c80-4337-9c81-54c67c23f2bc&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;improved HTC framework를 사용하는 52.3 box AP 및 46.0 mask AP의 상위 기준에서 Swin Transformer도 +4.1 box AP 및 +3.1 mask AP에서 높습니다(표 2(c) 참조).&lt;/li&gt;
&lt;/ul&gt;
&lt;figure id=&quot;044902e8-8832-429a-9a7c-167534d99491&quot; class=&quot;image&quot;&gt;&lt;a href=&quot;https://i.imgur.com/SYGWqyq.png[/img]&quot;&gt;&lt;img style=&quot;width: 466px;&quot; src=&quot;https://i.imgur.com/SYGWqyq.png[/img]&quot; /&gt;&lt;/a&gt;&lt;/figure&gt;
&lt;ul id=&quot;5ba112ac-209e-4317-8282-4e90bb041a64&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;추론 속도와 관련하여, ResNe(X)t는 고도로 최적화된 Cudnn 기능으로 구축된 반면, Swin-transformer 는 모두 최적화되지 않은 내장 PyTorch 기능으로 구현.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;36a2117f-5976-4cb7-bc94-39ed9b99dce7&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;kernel optimization는 본 논문의 범위를 벗어남.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li id=&quot;7740f6e0-f3a1-433a-b17e-7142d6705b89&quot; class=&quot;image&quot;&gt;&lt;a href=&quot;https://i.imgur.com/tQ2bDJZ.png[/img]&quot;&gt;&lt;img style=&quot;width: 477px;&quot; src=&quot;https://i.imgur.com/tQ2bDJZ.png[/img]&quot; /&gt;&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;fa51544b-d263-44a5-b646-851f15dd7eed&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Comparison to DeiT
&lt;ul id=&quot;b3f3fa83-1788-41b7-9eeb-0ec9706d2ce3&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;Cascade Mask R-CNN Framework를 이용한 DeiT-S의 성능을 표2(b)에 나타냄.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;3bbba53e-dd57-47c2-9d10-38019905e60c&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;Swin-T의 결과는 모델 크기가 비슷한 DeiT-S보다 +2.5 box AP와 +2.3 mask AP가 높고(86M vs 80M), 추론 속도도 상당히 빠름(15.3FPS vs 10.4FPS). DeiT의 추론 속도가 낮은 것은 주로 입력 영상 크기에 대한 quadratic complexity 때문.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;284afc1d-7a9a-4db7-aa84-a3cf13dc9d70&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Comparison to previous state-of-the-art
&lt;ul id=&quot;6dcc9235-0a78-4c12-b12c-ea1d9d2dc8eb&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;표 2(c)는 best results를 이전 state-ofthe-art models와 비교.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;ae0fdd8d-1d06-43b1-833e-899e17b53f59&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;논문의 베스트 모델은 COCO test-dev에서 58.7 box AP 및 51.1 mask AP를 달성하여 +2.7 box AP(외부 데이터가 없는 [25] Copy-paste) 및 +2.6 mask AP(DetectorRS [45])로 이전 최고의 결과를 능가.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;848f9a96-735f-49ec-a23e-10b35193b5f1&quot; data-ke-size=&quot;size23&quot;&gt;&amp;nbsp;Semantic Segmentation on ADE20K&lt;/h3&gt;
&lt;ul id=&quot;738ad5ba-4614-43ca-8390-34ec6787f8a0&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Settings
&lt;ul id=&quot;c3254f97-a20e-41cc-a7b6-5e999f265e35&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;ADE20K[80]는 널리 사용되는 semantic segmentation dataset으로, 150개의 semantic categories를 포괄. 총 25K개의 이미지를 보유하고 있으며, training 20K개, validation 2K개, testing 3K개.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;22fe6bcb-b480-45f8-8466-53576ccebad7&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;논문은 높은 효율성을 위한 기본 프레임워크로UperNet [66] in mmseg [15]을 활용.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;2211f028-b070-4578-8d2d-245c17ab397b&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;자세한 내용은 부록 참조.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;92d51674-81b8-47cd-b750-834bf0fd3373&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Results
&lt;figure id=&quot;01c8226b-0e8c-41b4-8208-73457d65e1d4&quot; class=&quot;image&quot;&gt;&lt;a href=&quot;https://i.imgur.com/skXBL75.png[/img]&quot;&gt;&lt;img style=&quot;width: 479px;&quot; src=&quot;https://i.imgur.com/skXBL75.png[/img]&quot; /&gt;&lt;/a&gt;&lt;/figure&gt;
&lt;ul id=&quot;17eacd98-e5a2-4569-9971-f50ce8100203&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;표 3은 different method/backbone pairs에 대한 mIoU, 모델 크기(#param), FLOP 및 FPS를 보여줌.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;00f4b347-09e4-4211-9358-db4a5152a76e&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;이 결과 비슷한 연산비용으로 Swin-S가 DeiT-S보다 +5.3mIoU(49.3 대 44.0) 높은 것으로 나타났다.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;e7af3e2a-c244-4d30-a090-8130ec40afa5&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;또한 ResNet-101보다 +4.4mIoU 높고 +2.4m입니다.ResNeSt-101[75]보다 높은 IoU.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;608be002-e148-4add-882a-17aac3c779f2&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;ImageNet-22K pre-training이 적용된 Swin-L 모델은 기존 최고 모델보다 +3.2mIoU(모델 크기가 더 큰 SETR [78])를 능가하는 53.5mIoU를 달성.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;35884829-f8e2-4786-95fa-87fd0da62627&quot; data-ke-size=&quot;size23&quot;&gt;&amp;nbsp;Ablation Study&lt;/h3&gt;
&lt;p id=&quot;633f5a0c-c116-4dec-b092-124b531a59e3&quot; class=&quot;&quot; data-ke-size=&quot;size16&quot;&gt;이 섹션에서는 ImageNet-1K image classification, COCO object detection Cascade Mask R-CNN, semantic segmentation 시 ADE20K UperNet을 사용하여 제안된 Swin Transformer에서 중요한 설계 요소를 단순화.&lt;/p&gt;
&lt;p id=&quot;06f4b8e7-bb1b-484d-b4d3-e6b0ee97a018&quot; class=&quot;&quot; data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;ul id=&quot;bd1ea438-5da4-431e-a71b-bb05091cc611&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Shifted windows
&lt;ul id=&quot;8fa294f7-5875-458e-a19c-28fc3236179d&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;세 가지 작업에 대한 Shifted window 접근법의 Ablations이 표 4에 보고.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;a8650b22-8a0b-40c1-98b6-7f4e3239d43b&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;Swin-T Shifted windows partitioning은 ImageNet-1K의 경우 +1.1% top-1 accuracy, COCO의 경우 +2.8 box AP/+2.2 mask AP, AD20K의 경우 +2.8 mioU만큼 각 단계에서 single window partitioning 은 우수.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;ebd1eb4b-32df-4c10-aeb5-59dc8694bc5c&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;결과는 preceding layers에서 Shifted windows을 사용하여 windows 간 연결을 구축하는 것의 효과를 나타냄.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;2c1d5a89-e2b2-4660-89ea-7e508e2b46cb&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;표 5와 같이 이동 창에 의한 latency overhead도 작습니다.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;28d158a8-4b82-40f4-9f90-078eebae6568&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Relative position bias
&lt;figure id=&quot;59e38ebd-f651-474a-b7db-a4953faf3cba&quot; class=&quot;image&quot;&gt;&lt;a href=&quot;https://i.imgur.com/MhfN6Da.png[/img]&quot;&gt;&lt;img style=&quot;width: 477px;&quot; src=&quot;https://i.imgur.com/MhfN6Da.png[/img]&quot; /&gt;&lt;/a&gt;&lt;/figure&gt;
&lt;ul id=&quot;bcf29a5d-7892-492d-a8ab-8a135cf7fe23&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;표 4는 position embedding approaches간의 차이를 보여줌. relative position bias가 있는 Swin-T는 ImageNet-1K에서 +1.2%/+0.8%의 top-1 accuracy, COCO에서 +1.3/+1.3 mask AP에서 +1.3/+1.3 mask AP에서 +2.9mioU 및 relation to those without position encoding and with absolute position embedding에서 +2.3/+2.9mioU를 산출.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;34d59551-4ce5-4cf2-8e23-1b5424946d72&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;또한 absolute position embedding을 포함하면 영상 분류 정확도(+0.4%)가 향상되지만, object detection and semantic segmentation(COCO의 경우 0.2 box/mask AP, ADE20K의 경우 -0.6mIoU)에 해를 미칩니다.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;7be28194-39b2-4ac6-b14d-a8f213e26ea2&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;최근image classification의 ViT/DeiT models의 abandon translation invariance은 오랫동안 시각적 모델링에 중요한 것으로 입증되었지만, 논문은 translation invariance을 장려하는 inductive bias이 general-purpose visual modeling, 특히 object detection and semantic segmentation의 dense prediction tasks에 여전히 선호된다는 것을 발견.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p id=&quot;e37570b3-f066-4c2e-a1cf-7a9f21d02a83&quot; class=&quot;&quot; data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;ul id=&quot;40f458a0-71c2-4eed-8547-e42942dcd472&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Different self-attention methods
&lt;figure id=&quot;0cba0e97-e2c0-471b-9966-4c3a0739dceb&quot; class=&quot;image&quot;&gt;&lt;a href=&quot;https://i.imgur.com/8qjSXOg.png[/img]&quot;&gt;&lt;img style=&quot;width: 400px;&quot; src=&quot;https://i.imgur.com/8qjSXOg.png[/img]&quot; /&gt;&lt;/a&gt;&lt;/figure&gt;
&lt;ul id=&quot;97194c7d-2242-4788-9b55-d026b5bc6231&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;다양한 self-attention computation과 구현의 실제 속도를 표 5에 비교.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;7dacf7b9-4cff-4684-b1ef-e4edfc7fb6ec&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;논문의 cyclic implementation은 특히 deeper stages에서 naive padding보다 하드웨어 효율성이 더 높음.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;d373d228-a1b9-4056-90ca-fb97ac465544&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;전체적으로 Swin-T, Swin-S, Swin-B에서 각각 13%, 18%의 속도 향상.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;9c688f15-73b3-4a5a-b7f5-0bd3a2ad8182&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;제안된 shifted windows 접근 방식을 기반으로 구축된 self-attention modules은 4개의 네트워크 단계에서 sliding windows보다 각각 40.8배/2.5배, 20.2배/2.5배, 9.3배/2.1배, 7.6배/1.8배 더 효율적.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;223901c9-de57-4e9a-bbcc-1083efe62ec2&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;전체적으로 shifted windows에 구축된 Swin Transformer 아키텍처는 Swin-T, Swin-S, Swin-B용 sliding windows에 구축된 변형 모델보다 각각 4.1/1.5, 4.0/1.5, 3.6/1.5배 더 빠름.&lt;/li&gt;
&lt;/ul&gt;
&lt;figure id=&quot;50a25a40-90fd-4a15-bd50-481981a5d2c4&quot; class=&quot;image&quot;&gt;&lt;a href=&quot;https://i.imgur.com/rJNuHdO.png[/img]&quot;&gt;&lt;img style=&quot;width: 480px;&quot; src=&quot;https://i.imgur.com/rJNuHdO.png[/img]&quot; /&gt;&lt;/a&gt;&lt;/figure&gt;
&lt;ul id=&quot;43c26eca-3983-4753-a1d1-07b1a5656823&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;표 6은 세 가지 작업에 대한 정확성을 비교하여 시각적 모델링에 있어 similarly accurate을 보여줌.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;d11b1a63-7878-4b59-a908-a6a30916c89c&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;가장 빠른 트랜스포머 아키텍처 중 하나인 Performer [13]에 비해([59] 참조), 제안된 shifted windows attention computation 과 전체 Swin Transformer 아키텍처는 약간 빠르며(표 5 참조), Swin-T를 사용하는 ImageNet-1K에 비해 +2.3%의 상위 1위 정확도를 달성(표 6 참조).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;a93be555-5b1b-4948-956d-74cc9196e16d&quot; data-ke-size=&quot;size26&quot;&gt;Conclusion&lt;/h2&gt;
&lt;ul id=&quot;89f4be29-57ea-43ef-9ef7-bee70ed2ca2b&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;이 논문에서는 hierarchical feature representation을 생산하고 입력 이미지 크기에 대한 linear computational complexity을 갖는 새로운 비전 Transformer인 Swin Transformer를 소개.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;dbfbb61a-6389-42ae-bd4a-a6634a810632&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Swin Transformer는 COCO Object detection와 관련하여 SOTA를 달성.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;7be3aa10-d834-4513-925b-6bb896bdb54a&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;ADE20K semantic segmentation, 이전 SOTA를 훨씬 능가.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;c34cc7c6-33a9-43c8-85fc-48d5314279cf&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;다양한 비전 문제에 대한 Swin Transformer의 강력한 성능이 vision and language signals의 통일된 모델링을 촉진하기를 바람.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;016fff8f-794d-4ff1-91fc-2ade01f85303&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Swin Transformer의 핵심 요소로서, shifted window based self-attention이 비전 문제에 효과적이고 효율적인 것으로 입증되었으며, 자연어 처리에서도 활용도를 조사할 수 있기를 기대.&lt;/li&gt;
&lt;/ul&gt;
&lt;p id=&quot;ef0230f7-966f-4ab3-9321-2a70d94ebef1&quot; class=&quot;&quot; data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h2 id=&quot;21f9a522-c916-41d4-9d11-556fe5a3c4b9&quot; data-ke-size=&quot;size26&quot;&gt;A1. Detailed Architectures&lt;/h2&gt;
&lt;figure id=&quot;28aaf657-620b-48b7-a60d-e7286d48cd5f&quot; class=&quot;image&quot;&gt;&lt;a href=&quot;https://i.imgur.com/FXsVgKd.png[/img]&quot;&gt;&lt;img style=&quot;width: 768px;&quot; src=&quot;https://i.imgur.com/FXsVgKd.png[/img]&quot; /&gt;&lt;/a&gt;&lt;/figure&gt;
&lt;ul id=&quot;0382e610-176e-4cae-8e0a-a29e1a0592fd&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;자세한 아키텍처 사양은 모든 아키텍처에 대해 224&amp;times;224의 입력 이미지 크기를 가정하는 표 7에 나와 있음.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;2fb8f9dd-b6a2-41e5-afb7-322a12cacfff&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;&amp;ldquo;Concat $n&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt; \times n$ &lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;은 패치에서 $&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;n&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mbin&quot;&gt;&amp;times;&lt;/span&gt;&lt;span style=&quot;margin-right: 0.2222222222222222em;&quot; class=&quot;mspace&quot;&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.43056em; vertical-align: 0em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord mathnormal&quot;&gt;n$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;이웃 피쳐의 연결을 나타냄. 이 작업을 수행하면 feature map의 다운샘플링 속도가 n.
&lt;ul id=&quot;c2d3a55a-e8dc-410a-aee2-c8640c72868b&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;&quot;96-d&quot;는 출력 dim이 96인 Linear layer를 나타냄. &quot;win. sz. 7 &amp;times; 7&quot;은 window size가 7 &amp;times; 7인 multi-head self-attention (MSA) module을 나타냄.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;b3baafa9-b8da-40f3-bbb7-e3ec703f9b7a&quot; data-ke-size=&quot;size26&quot;&gt;A2. Detailed Experimental Settings&lt;/h2&gt;
&lt;h3 id=&quot;5100ea89-f217-466a-8829-4263c287a8a3&quot; class=&quot;&quot; data-ke-size=&quot;size23&quot;&gt;A2.1. Image classiﬁcation on ImageNet-1K&lt;/h3&gt;
&lt;ul id=&quot;2af8508d-98ae-4928-8b6a-52d1e878d839&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;image classification는 마지막 단계의 출력 피쳐 맵에 global average pooling layer를 적용한 다음 linear classifier를 적용하여 수행.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;c60f1d4f-86fe-45a9-b087-144e75744573&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;이 전략이 ViT[19]와 DeiT[60]에서와 같이 additional class 토큰을 사용하는 것만큼 정확하다고 생각함.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;1652c75b-5248-40ee-845e-65a526d30ecf&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;평가에서 single crop를 사용한 top-1 accuracy 가 보고.&lt;/li&gt;
&lt;/ul&gt;
&lt;p id=&quot;100588bc-c13b-4c66-901c-f86aec3b01e8&quot; class=&quot;&quot; data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;ul id=&quot;08afb9a9-671a-4ef1-96ff-e0af6e2cd6c4&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Regular ImageNet-1K training
&lt;ul id=&quot;cdb0762b-9a06-4ce1-9886-be0355a41ba9&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;대부분의 training settings은 [60]을 따름.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;669d2678-ecf2-4141-92a2-ba6036b27c2a&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;모든 모델 변형에 대해 기본 입력 이미지 해상도$224^2$ 를 채택.$&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;384^2$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;와 같은 다른 해상도의 경우 GPU 소비를 줄이기 위해 처음부터 교육하는 대신 $&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;224^2$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&amp;nbsp;해상도로 교육된 모델을 fine-tune.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;54864523-de56-42e3-9c2d-e3a6e787a42b&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;$224^2$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;input 으로 처음부터 training할 때, 20개의 epochs of linear warm-up이 있는 cosine decay learning rate scheduler를 사용하여 300개 epoch에 AdamW[36] 최적화 사용.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;5a01e879-cf89-4afc-950a-d1b48779b724&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;batch size 1024, initial learning rate 0.001, weight decay 0.05 및 max norm 1의 gradient clipping이 사용됨.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;1facd9e9-1082-4c05-baed-872c3dd60860&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;RandAugment [16], Mixup [74], Cutmix [72], random erasing [79], stochastic depth [34]를 포함하여 대부분의 augmentation 및 정규화 전략을 training에 포함하지만 repeated augmentation[30] 및 Exponential Moving Average[44]은 포함하지 않음.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;0c6ef4bc-d919-473d-b81c-0a606ec44026&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;이는 ViT 훈련을 안정화하기 위해 repeated augmentation가 중요하다는 점과 반대되는 것에 유의.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;f5697a2b-23e1-4e44-8cca-883c2c486031&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;대형 모델(예: 각각 0.2; 0.3; 0.5), SWin-T, SWin-S 및 SWin-B의 경우 0.5)에 stochastic depth augmentation가 사용된다.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;5e67b7f2-77aa-48a0-8f27-e83c2d48b326&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;분해능이 더 큰 입력에 대한 미세 조정을 위해, 확률적 깊이 비율을 0.1로 설정하는 것을 제외하고, 10-5의 일정한 학습 속도, 10-8의 weight decay, 첫 번째 단계와 동일한 데이터 augmentation 및 정규화의 30개 에포크에 대해 Adam W[36] 최적화 장치를 사용한다.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;2a1cea6f-1fe5-4755-9638-1d16475447d8&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;ImageNet-22K pre-training
&lt;ul id=&quot;d07a77e8-fe85-4186-9178-6330a5197c21&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;또한 14.2 million 개의 이미지와 22K 클래스를 포함하는 대규모 ImageNet-22K 데이터셋에 대해 pre-training.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;aff4a98c-b0e2-4ce3-8327-df2ed7b37de7&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: circle;&quot;&gt;training은 두 단계로 진행.
&lt;ul id=&quot;1b3155b6-4d84-486d-bc25-de67f36e2938&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;$224^2$입력의 첫 번째 단계에서는 5 epochs linear warm-up을 사용하는 linear warm-up scheduler를 사용하여 60epochs에 AdamW optimizer를 사용.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;d55114f9-7973-4092-a140-7c160fe38709&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;배치 크기 4096, 초기 학습률 0.001 및 weight decay 0.01이 사용.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;f9c77a85-7c47-4f45-a899-8bde0c8a1b19&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: square;&quot;&gt;&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;$224^2/384^2$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&amp;nbsp;입력으로 ImageNet-1K finetuning의 두 번째 단계에서는 배치 크기 1024, 일정한 학습 속도 10^-5, weight decay 10^-8의 30epoch 모델을 training.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;581cf740-480f-42ac-95f7-621d29fb471b&quot; class=&quot;&quot; data-ke-size=&quot;size23&quot;&gt;A2.2. Object detection on COCO&lt;/h3&gt;
&lt;ul id=&quot;646f28bb-364a-4793-a27b-cc983a5da656&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;ablation study의 경우, 네 가지 일반적인 개체 탐지 프레임워크인 Cascade Mask R-CNN [28, 5], ATSS [76], RepPoints v2 [11], and Sparse RCNN [55] in mmdetection [9]을 고려.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;4cf85807-ca59-483a-8135-2b8e1490a2d8&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;이 네 가지 프레임워크에 대해, 우리는 동일한 설정을 활용:multi-scale training [7, 55] (짧은 쪽이 480에서 800 사이인 반면 긴 쪽이 최대 1333 사이인 입력의 크기 조정), Adam W[43] optimizer(초기 학습률 0.0001, weight decay 0.05, 배치 크기 16), 3x schedule(learning rate decayed 36 epochs 27과 33에서 10배 증가).&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;25326e10-2194-4284-8280-b5d3831ec704&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;시스템 수준 비교를 위해, 우리는 향상된HTC[8] (HTC++로 표시), instaboost [21], 보다 강력한 multi-scale training[6] (짧은 쪽이 400~1400 사이인 반면 긴 쪽이 최대 1600까지), 6x schedule(72 epochs(63~69에 학습 속도가 0.1배 감소), soft-NMS [4], 마지막 단계의 출력과 ImageNet-22K pre-trained 모델을 초기화.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;68307f6e-720c-4389-90dc-6f3c7d9bd408&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;모든 Swin Transformer 모델에 대해 0.2의 비율로 stochastic depth를 채택&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;e209c3da-0a65-4e0f-a329-47de4652b5cc&quot; class=&quot;&quot; data-ke-size=&quot;size23&quot;&gt;A2.3. Semantic segmentation on ADE20K&lt;/h3&gt;
&lt;ul id=&quot;f98a156c-a921-4f49-b8b9-6436d5d920ac&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;ADE20K[80]는 널리 사용되는 semantic segmentation dataset으로, 150개의 semantic categories를 포괄&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;6239ab00-9eff-4871-bf0d-40e4113d7fe7&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;총 25K개의 이미지를 보유하고 있으며, training 20K개, validation 2K개, testing 3K개.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;6ad817dd-c4d0-476d-ad79-24168ece1ed3&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;높은 효율성을 위한 기본 프레임워크로 UperNet[66] in mm segmentation[15]을 활용합니다.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;ef0d7246-af89-428d-bd6d-90868e749b22&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;training 에서는 initial learning rate이 $6 &amp;times; 10-5$, weight decay가 0.01, linear learning rate decay를 사용하는 scheduler 및 1,500회 반복의 linear warmup을 사용하는 AdamW[43] 최적화기를 사용.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;b786631d-5232-40e6-aca9-38d9866d0115&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;모델은 GPU당 이미지 2개가 포함된 8개의 GPU에서 160K회 반복 training 을 받음.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;f300eb23-d942-43cf-9f51-6b9a589c4f47&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;augmentations의 경우, random horizontal flipping의 mmsegmentation, [0.5, 2.0]비율 범위 내 random re-scaling 및 random photometric distortion의 기본 설정을 채택.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;0fd0e0ab-2036-4084-b0ba-8f861a096afe&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;모든 SwinTransformer 모델에 0.2의 Stochastic depth가 적용.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;b40cff37-5b30-4a74-bf07-4a61e41dfd2c&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;Swin-T, Swin-S는 512&amp;times;512의 입력으로 이전 접근 방식에 따라 standard setting에 대한 trained.&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;&amp;Dagger;&amp;nbsp;\ddagger&amp;nbsp;&lt;/span&gt;&lt;span class=&quot;katex-html&quot; aria-hidden=&quot;true&quot;&gt;&lt;span class=&quot;base&quot;&gt;&lt;span style=&quot;height: 0.8888799999999999em; vertical-align: -0.19444em;&quot; class=&quot;strut&quot;&gt;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&amp;Dagger;&lt;/span&gt;&lt;span class=&quot;mord&quot;&gt;&amp;nbsp;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;가 있는 Swin-B 및 Swin-L은 이 두 모델이 ImageNet-22K에서 pre-trained되었으며 640&amp;times;640의 입력으로 trained되었음을 나타냄.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;c2530dae-783f-4f70-a7d2-dcc899ed49a0&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;inference에서는 training 에 사용되는 resolutions의 $[0.5, 0.75, 1.0, 1.25, 1.5, 1.75]&amp;times;$을 사용한 multi-scale test가 사용.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;f4585718-f0ff-4791-9b2a-26f5be0cbfc3&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;test scores를 보고할 때, common practice에 따라 training images and validation images이 모두 training 에 사용[68].&lt;/li&gt;
&lt;/ul&gt;
&lt;p id=&quot;0dc2a37a-5bc8-43e1-9871-26b770cef6b6&quot; class=&quot;&quot; data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h2 id=&quot;ffd3a6bc-791f-4d0e-9b9b-e1aa6081ca1d&quot; data-ke-size=&quot;size26&quot;&gt;A3. More Experiments&lt;/h2&gt;
&lt;h3 id=&quot;54142cbd-90f7-4bc4-a542-24076012c322&quot; class=&quot;&quot; data-ke-size=&quot;size23&quot;&gt;A3.1. Image classiﬁcation with different input size&lt;/h3&gt;
&lt;figure id=&quot;7e2d2c0b-69f9-4468-8832-03663e764a6f&quot; class=&quot;image&quot;&gt;&lt;a href=&quot;https://i.imgur.com/R0b1KrU.png[/img]&quot;&gt;&lt;img style=&quot;width: 485px;&quot; src=&quot;https://i.imgur.com/R0b1KrU.png[/img]&quot; /&gt;&lt;/a&gt;&lt;/figure&gt;
&lt;ul id=&quot;dccf791c-5f74-4fa3-bde9-67fe31bc23ed&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;표 8은$&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;224^2$&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;부터 $&lt;span style=&quot;user-select: all; -webkit-user-select: all; -moz-user-select: all;&quot; class=&quot;notion-text-equation-token&quot; data-token-index=&quot;0&quot;&gt;&lt;span&gt;&lt;span class=&quot;katex&quot;&gt;&lt;span class=&quot;katex-mathml&quot;&gt;384^2$ &lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;까지의 다양한 입력 이미지 크기를 가진 Swin Transformer의 성능을 보여줍니다. 일반적으로 입력 resolution이 클수록 top-1 accuracy는 향상되지만 추론 속도는 느려짐.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;b0357deb-afc9-460f-a07e-a594681ff807&quot; class=&quot;&quot; data-ke-size=&quot;size23&quot;&gt;A3.2. Different Optimizers for ResNe(X)t on COCO&lt;/h3&gt;
&lt;figure id=&quot;3caad3ec-7e8b-4e45-8ae4-878a84cb1886&quot; class=&quot;image&quot;&gt;&lt;a href=&quot;https://i.imgur.com/PzZ94eD.png[/img]&quot;&gt;&lt;img style=&quot;width: 475px;&quot; src=&quot;https://i.imgur.com/PzZ94eD.png[/img]&quot; /&gt;&lt;/a&gt;&lt;/figure&gt;
&lt;ul id=&quot;ebd93ed0-0d39-4560-895d-5564127a3cac&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;표 9는 COCO 객체 감지 시 ResNe(X)t 백본의 AdamW 및 SGD 최적화기를 비교.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;cd691ab2-dedb-468d-af60-4b21214bf92b&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;이 비교에는 Cascade Mask R-CNN 프레임워크가 사용. SGD는 Cas-cade Mask R-CNN 프레임워크의 기본 optimizer로 사용되지만, 일반적으로 특히 작은 백본의 경우 AdamW optimizer로 교체하여 정확도가 향상되는 것을 관찰.&lt;/li&gt;
&lt;/ul&gt;
&lt;ul id=&quot;c8423d6a-d6c8-4b4d-8806-67196ea7e167&quot; class=&quot;bulleted-list&quot; style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li style=&quot;list-style-type: disc;&quot;&gt;따라서 제안된 Swin Transformer 아키텍처와 비교할 때 AdamW for ResNe(X)t 백본을 사용.&lt;/li&gt;
&lt;/ul&gt;
&lt;p id=&quot;8a9e1062-402c-4ca8-915b-d8b7a8d7426e&quot; class=&quot;&quot; data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p id=&quot;0e7dc362-6730-47ba-b4d9-f17c27b84d26&quot; class=&quot;&quot; data-ke-size=&quot;size16&quot;&gt;Q. 기존 선행논문인 ViT에 있는 [CLS] Token 은 어디?&lt;/p&gt;
&lt;/div&gt;
&lt;/article&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;</description>
      <category>인공지능/Computer Vision</category>
      <author>통경</author>
      <guid isPermaLink="true">https://stat-cbc.tistory.com/37</guid>
      <comments>https://stat-cbc.tistory.com/37#entry37comment</comments>
      <pubDate>Wed, 22 Sep 2021 23:32:45 +0900</pubDate>
    </item>
    <item>
      <title>DANet(Dual Attention Network for Scene Segmentation) 논문 리뷰 - CVPR_2019</title>
      <link>https://stat-cbc.tistory.com/30</link>
      <description>&lt;h3 data-ke-size=&quot;size23&quot;&gt;&lt;span&gt;&amp;nbsp;Abstract&lt;/span&gt;&lt;/h3&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;Self Attention Mechanism 을 기반으로 다양한 상황 의존성을 캡처하여 Scene Segmentation 수행&lt;/li&gt;
&lt;li&gt;multi-scale feature fusion 으로 context를 포착하는 이전 논문(ICEnet, ... ) 과 달리, local feature 를 global dependencies과 적응적(adaptively)으로 통합할 수 있는 DANet (Dual Attention Network)을 제안 (position + channel Attention)&lt;/li&gt;
&lt;li&gt;두 가지 유형의 Attention 모듈을 dilated FCN 모듈&amp;nbsp;위에 추가&lt;/li&gt;
&lt;li&gt;position 및 channel 에서 각각 semantic interdependencies을 모델링&lt;/li&gt;
&lt;/ul&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;1) position Attention module : &lt;/span&gt;&lt;span&gt;모든 &lt;/span&gt;&lt;span&gt;position&lt;/span&gt;&lt;span&gt;에서 &lt;/span&gt;&lt;span&gt;feature &lt;/span&gt;&lt;span&gt;의 가중치 합계를 사용하여 각 &lt;/span&gt;&lt;span&gt;position&lt;/span&gt;&lt;span&gt;의 &lt;/span&gt;&lt;span&gt;feature &lt;/span&gt;&lt;span&gt;를 선택적으로 집계&lt;/span&gt;&lt;span&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;- &lt;/span&gt;&lt;span&gt;유사한 특징은 거리에 관계 없이 서로 연관&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;2) channel Attention module : &lt;/span&gt;&lt;span&gt;모든 &lt;/span&gt;&lt;span&gt;channel map &lt;/span&gt;&lt;span&gt;사이에 관련 &lt;/span&gt;&lt;span&gt;feature &lt;/span&gt;&lt;span&gt;을 통합하여 상호 의존적인 &lt;/span&gt;&lt;span&gt;channel map&lt;/span&gt;&lt;span&gt;을 선택적으로 강조&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;&amp;gt; &lt;/span&gt;&lt;span&gt;두 &lt;/span&gt;&lt;span&gt;Attention &lt;/span&gt;&lt;span&gt;모듈의 출력을 합산하여 &lt;/span&gt;&lt;span&gt;feature representation&lt;/span&gt;&lt;span&gt;을 개선하여 보다 정확한 &lt;/span&gt;&lt;span&gt;Segmentataion &lt;/span&gt;&lt;span&gt;결과에 기여 &lt;/span&gt;&lt;/p&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;세 종류의 Scene Segmentation 데이터 세트(Cityscapes, PASCAL Context 및 COCO Stuff)에서 SOTA 성능 달성.&lt;/li&gt;
&lt;li&gt;특히 Cityscapes 테스트셋에서 coarse 데이터를 사용하지 않고도 81.5% 의 Mean IoU 를 얻음&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;&lt;span&gt;Introduction&lt;/span&gt;&lt;/h3&gt;
&lt;p data-ke-size=&quot;size18&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;- Scene Segmentation &lt;/span&gt;&lt;span&gt;작업을 효과적으로 수행하기 위해서는 혼란스러운 범주를 구별하고 외관이 다른 객체를 고려해야 함&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;예&lt;/span&gt;&lt;span&gt;) '&lt;/span&gt;&lt;span&gt;밭&lt;/span&gt;&lt;span&gt;'&lt;/span&gt;&lt;span&gt;과 &lt;/span&gt;&lt;span&gt;'&lt;/span&gt;&lt;span&gt;풀&lt;/span&gt;&lt;span&gt;'&lt;/span&gt;&lt;span&gt;의 영역은 종종 구별할 수 없으며&lt;/span&gt;&lt;span&gt;, '&lt;/span&gt;&lt;span&gt;자동차&lt;/span&gt;&lt;span&gt;'&lt;/span&gt;&lt;span&gt;의 물체는 &lt;/span&gt;&lt;span&gt;scales, occlusion, illumination&lt;/span&gt;&lt;span&gt;에 의해 영향을 받음&lt;/span&gt;&lt;span&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;&amp;nbsp;픽셀 수준 인식을 통해 feature 표현의 discriminative ability을 강화할 필요성이 있음.&lt;/li&gt;
&lt;li&gt;Fully Convolutional Networks (FCNs)[13] 모델이 state-of-the-art 성능 보임.&amp;nbsp;&lt;/li&gt;
&lt;/ul&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;- &lt;/span&gt;&lt;span&gt;선행연구&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;1) multi-scale context fusion &lt;/span&gt;&lt;span&gt;활용 &lt;/span&gt;&lt;span&gt;: Deeplab, PSPNet... &lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;2) long-range dependencies &lt;/span&gt;&lt;span&gt;모형화를 위한 &lt;/span&gt;&lt;span&gt;recurrent neural network : 2D-LSTM(local features&lt;/span&gt;&lt;span&gt;에 대한 풍부한 &lt;/span&gt;&lt;span&gt;spatial dependencies&lt;/span&gt;&lt;span&gt;을 포착하기 위해 &lt;/span&gt;&lt;span&gt;directed acyclic graph&lt;/span&gt;&lt;span&gt;를 사용하여 &lt;/span&gt;&lt;span&gt;recurrent neural networks&lt;/span&gt;&lt;span&gt;을 구축&lt;/span&gt;&lt;span&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;natural scene image 분할을 위한 새로운 프레임워크 DANet(Dual Attention Network) 제안(그림 2 참조).&lt;/li&gt;
&lt;li&gt;spatial and channel dimensions의 각각 feature dependencies 을 포착하기 위한 self-Attention mechanism을 도입.&lt;/li&gt;
&lt;li&gt;dilated FCN 위에 두 개의 병렬 Attention 모듈을 추가.&lt;/li&gt;
&lt;/ul&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;- 구조&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;natural scene image &lt;/span&gt;&lt;span&gt;분할을 위한 새로운 프레임워크 &lt;/span&gt;&lt;span&gt;DANet(Dual Attention Network) &lt;/span&gt;&lt;span&gt;제안&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;그림 &lt;/span&gt;&lt;span&gt;2 &lt;/span&gt;&lt;span&gt;참조&lt;/span&gt;&lt;span&gt;).&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;spatial and channel dimensions&lt;/span&gt;&lt;span&gt;의 각각 &lt;/span&gt;&lt;span&gt;feature dependencies &lt;/span&gt;&lt;span&gt;을 포착하기 위한 &lt;/span&gt;&lt;span&gt;self-Attention mechanism&lt;/span&gt;&lt;span&gt;을 도입&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;dilated FCN &lt;/span&gt;&lt;span&gt;위에 두 개의 병렬 &lt;/span&gt;&lt;span&gt;Attention &lt;/span&gt;&lt;span&gt;모듈을 추가&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;1) position Attention module : feature map&lt;/span&gt;&lt;span&gt;의 두 &lt;/span&gt;&lt;span&gt;position &lt;/span&gt;&lt;span&gt;사이의 공간 의존성을 캡처하기 위&lt;/span&gt;&lt;span&gt;한 &lt;/span&gt;&lt;span&gt;self-Attention mechanism.&amp;nbsp;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;가중치는 해당 두 위치 사이의 유사도&lt;/span&gt;&lt;span&gt;.&amp;nbsp;&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;두 &lt;/span&gt;&lt;span&gt;position &lt;/span&gt;&lt;span&gt;간 거리에 유사도는 관계없음&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;2) channel Attention module : Self-Attention mechanism&lt;/span&gt;&lt;span&gt;을 사용하여 두 채널 맵 사이&lt;/span&gt;&lt;span&gt;의 종속성을 캡처&lt;/span&gt;&lt;span&gt;.&amp;nbsp;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;3) &lt;/span&gt;&lt;span&gt;이 두 &lt;/span&gt;&lt;span&gt;Attention &lt;/span&gt;&lt;span&gt;모듈의 출력은 &lt;/span&gt;&lt;span&gt;feature &lt;/span&gt;&lt;span&gt;표현을 더욱 강화하기 위해 &lt;/span&gt;&lt;b&gt;&lt;span&gt;Sum fusion.&lt;/span&gt;&lt;/b&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;복잡하고 다양한 장면을 다룰 때 이전 방법&lt;/span&gt;&lt;span&gt;[4, 29]&lt;/span&gt;&lt;span&gt;보다 더 효과적이고 유연&lt;/span&gt;&lt;span&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;388&quot; data-origin-height=&quot;305&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/oKvp4/btrdSeS5Hj4/k1fWPxiKsemsyJLs8yqk5K/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/oKvp4/btrdSeS5Hj4/k1fWPxiKsemsyJLs8yqk5K/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/oKvp4/btrdSeS5Hj4/k1fWPxiKsemsyJLs8yqk5K/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FoKvp4%2FbtrdSeS5Hj4%2Fk1fWPxiKsemsyJLs8yqk5K%2Fimg.png&quot; data-origin-width=&quot;388&quot; data-origin-height=&quot;305&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;예)&amp;nbsp;그림1&amp;nbsp;의 길거리 장면.&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;1) &lt;/span&gt;&lt;span&gt;첫 번째 사진의 일부 &lt;/span&gt;&lt;span&gt;'&lt;/span&gt;&lt;span&gt;사람&lt;/span&gt;&lt;span&gt;'&lt;/span&gt;&lt;span&gt;과 &lt;/span&gt;&lt;span&gt;'&lt;/span&gt;&lt;span&gt;신호등&lt;/span&gt;&lt;span&gt;'&lt;/span&gt;&lt;span&gt;은 조명과 시야로 인해 눈에 띄지 않거나 불완전한 물체&lt;/span&gt;&lt;span&gt;. &lt;/span&gt;&lt;span&gt;큰 물체&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;예&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;자동차&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;건물&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;span&gt;의 맥락이 눈에 띄지 않는 물체 라벨링을 해칠 수 있음&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;반대로&lt;/span&gt;&lt;span&gt;, Attention &lt;/span&gt;&lt;span&gt;모델은 눈에 띄지 않는 객체의 유사한 특징을 선별적으로 취합하여 특징 표현을 강조하고 큰 물체의 영향을 피함&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;2) '&lt;/span&gt;&lt;span&gt;자동차&lt;/span&gt;&lt;span&gt;'&lt;/span&gt;&lt;span&gt;와 &lt;/span&gt;&lt;span&gt;'&lt;/span&gt;&lt;span&gt;사람&lt;/span&gt;&lt;span&gt;'&lt;/span&gt;&lt;span&gt;의 &lt;/span&gt;&lt;span&gt;scale&lt;/span&gt;&lt;span&gt;은 다양하며&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;그러한 다양한 대상을 인식하기 위해서는 다른 &lt;/span&gt;&lt;span&gt;scale &lt;/span&gt;&lt;span&gt;의 상황 정보가 필요&lt;/span&gt;&lt;span&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;즉&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;서로 다른 &lt;/span&gt;&lt;span&gt;scale&lt;/span&gt;&lt;span&gt;의 &lt;/span&gt;&lt;span&gt;feature&lt;/span&gt;&lt;span&gt;들은 동일한 &lt;/span&gt;&lt;span&gt;semantic&lt;/span&gt;&lt;span&gt;를 나타내기 위해 동등하게 다루어져야 함&lt;/span&gt;&lt;span&gt;. Attention mechanism&lt;/span&gt;&lt;span&gt;이 있는 우리 모델은 &lt;/span&gt;&lt;span&gt;global view&lt;/span&gt;&lt;span&gt;에서 어떤 규모로든 유사한 &lt;/span&gt;&lt;span&gt;feature &lt;/span&gt;&lt;span&gt;을 적응적으로 통합하는 것을 목표&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;3) position &lt;/span&gt;&lt;span&gt;과 &lt;/span&gt;&lt;span&gt;channel&lt;/span&gt;&lt;span&gt;관계를 &lt;/span&gt;&lt;span&gt;explicitly&lt;/span&gt;&lt;span&gt;하게고려 &lt;/span&gt;&lt;span&gt;long-range dependencies&lt;/span&gt;&lt;span&gt;에 유익&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;&lt;span&gt;Relate work&lt;/span&gt;&lt;/h3&gt;
&lt;p data-ke-size=&quot;size18&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;1) multi-scale feature fusion &lt;/span&gt;&lt;/p&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;Deeplabv2 [3]와 Deeplabv3 [4] : Contextual information&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;를 포함하기 위해 &lt;/span&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;atrous spatial pyramid pooling&lt;/span&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;을 채택&lt;/span&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;, &lt;/span&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;이는 서로 다른 &lt;/span&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;dilated &lt;/span&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;속도를 가진 &lt;/span&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;parallel dilated &lt;/span&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;컨볼루션으로 구성&lt;/span&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;.&lt;/span&gt;&lt;/li&gt;
&lt;li&gt;PSP-Net [29] : 서로 다른 scales정보를 포함하는 효과적인 Contextual prior를 위한 pyramid pooling module.&lt;/li&gt;
&lt;li&gt;인코더-디코더 구조[6, 8, 9] : 다른 scale 컨텍스트를 얻기 위해 중간 레벨 및 높은 레벨의 semantic feature 을 결합.&lt;/li&gt;
&lt;/ul&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;2) &lt;/span&gt;&lt;span&gt;로컬 &lt;/span&gt;&lt;span&gt;feature&lt;/span&gt;&lt;span&gt;에 대한 &lt;/span&gt;&lt;span&gt;Contextual &lt;/span&gt;&lt;span&gt;의존성을 학습하는 것도 &lt;/span&gt;&lt;span&gt;feature &lt;/span&gt;&lt;span&gt;표현에 기여&lt;/span&gt;&lt;span&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;DAG-RNN [18] : &lt;span style=&quot;letter-spacing: 0px;&quot;&gt;풍부한 &lt;/span&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;Contextual dependencies&lt;/span&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;을 포착하기 위해 &lt;/span&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;Recurrent neural network&lt;/span&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;을 가진 &lt;/span&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;directed acyclic graph &lt;/span&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;사용&lt;/span&gt;&lt;span style=&quot;letter-spacing: 0px;&quot;&gt;.&lt;/span&gt;&lt;/li&gt;
&lt;li&gt;PSANet [30] :&amp;nbsp;spatial dimension의 relative position information과 convolution layer를 기준으로 pixel-wise relation를 captures.&lt;/li&gt;
&lt;li&gt;EncNet [27] : global context를 포착하기 위한 channel Attention mechanism을 도입.&lt;/li&gt;
&lt;/ul&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;3) selfAttention &lt;/span&gt;&lt;span&gt;모듈은 장거리 의존성을 모델링할 수 있어 광범위하게 적용 &lt;/span&gt;&lt;span&gt;[11, 12, 17, 19&lt;/span&gt;&lt;span&gt;&amp;ndash;&lt;/span&gt;&lt;span&gt;21].&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;[21] :&amp;nbsp; 입력의 global dependencies을 도출하기 위한 self-Attention mechanism을 제안하는 첫 번째 작업이며 기계 번역에 이를 적용.&lt;/li&gt;
&lt;li&gt;[28] : Attention Module은 이미지에 점점 더 많이 적용. 더 나은 image generator를 학습하기 위한 self Attention mechanism을 도입.&lt;/li&gt;
&lt;li&gt;[23] : 주로 영상과 이미지에 대한 시공간 차원에서 non-local 효과를 탐구.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;&lt;span&gt;Dual Attention Network&lt;/span&gt;&lt;/h3&gt;
&lt;p data-ke-size=&quot;size18&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;네트워크의 일반적인 프레임워크를 제시&lt;/span&gt;&lt;span&gt;. position &lt;/span&gt;&lt;span&gt;및 &lt;/span&gt;&lt;span&gt;channel &lt;/span&gt;&lt;span&gt;에서 각각 장거리 상황 정보를 캡처하는 두 가지 &lt;/span&gt;&lt;span&gt;Attention &lt;/span&gt;&lt;span&gt;모듈을 소개&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;h4 data-ke-size=&quot;size20&quot;&gt;&lt;span&gt;&amp;lt; Overview&amp;gt;&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;799&quot; data-origin-height=&quot;466&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/bC3V3a/btrdOMv1duc/8xID2YdZlSkj76kzIBK8R1/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/bC3V3a/btrdOMv1duc/8xID2YdZlSkj76kzIBK8R1/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/bC3V3a/btrdOMv1duc/8xID2YdZlSkj76kzIBK8R1/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FbC3V3a%2FbtrdOMv1duc%2F8xID2YdZlSkj76kzIBK8R1%2Fimg.png&quot; data-origin-width=&quot;799&quot; data-origin-height=&quot;466&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;그림 &lt;/span&gt;&lt;span&gt;2. dilated residual network&lt;/span&gt;&lt;span&gt;에 의해 생성된 &lt;/span&gt;&lt;span&gt;local feature&lt;/span&gt;&lt;span&gt;를 통해 &lt;/span&gt;&lt;span&gt;global context&lt;/span&gt;&lt;span&gt;를 위해 두 가지 유형의 &lt;/span&gt;&lt;span&gt;Attention &lt;/span&gt;&lt;span&gt;모듈으로 더 나은 &lt;/span&gt;&lt;span&gt;feature &lt;/span&gt;&lt;span&gt;표현을 얻고자 함&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h4 data-ke-size=&quot;size20&quot;&gt;&lt;span&gt;&amp;lt; Position Attention Module&amp;gt;&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;508&quot; data-origin-height=&quot;223&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/cw1OU9/btrdPJTzqph/sgYELcutoaMHx1Ekk7rXk0/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/cw1OU9/btrdPJTzqph/sgYELcutoaMHx1Ekk7rXk0/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/cw1OU9/btrdPJTzqph/sgYELcutoaMHx1Ekk7rXk0/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2Fcw1OU9%2FbtrdPJTzqph%2FsgYELcutoaMHx1Ekk7rXk0%2Fimg.png&quot; data-origin-width=&quot;508&quot; data-origin-height=&quot;223&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;Discriminant feature representations은 scene understanding 를 위해 필수적, 이는 long-range contextual information를 얻을 수 있음.&lt;/li&gt;
&lt;li&gt;그러나 많은 연구[15, 29]에서 기존 FCN에서 생성된 local features 가 객체를 잘못 분류할 수 있다고 주장. local features 에 대한 풍부한 컨텍스트 관계를 모델링하기 위해 position attention module 도입.&lt;/li&gt;
&lt;li&gt;position Attention module 은 보다 광범위한 상황 정보를 local feature으로 인코딩하여 표현 능력을 향상시킨 후 적응적으로 spatial 컨텍스트를 집계하는 프로세스를 자세히 설명&lt;/li&gt;
&lt;li&gt;그림.3(A)에 예시된 바와 같이, local feature $ \in ℝ^{ &amp;times;H&amp;times;W}$가 주어졌을 때, 두 개의 새로운 feature maps B와 C를 생성하기 위해 convolution layer 적용 $B \in ℝ^{ &amp;times;N}$으로 reshape. $N = H \times W$&lt;/li&gt;
&lt;li&gt;그런 다음 C와 B의 전치 사이에 행렬 곱셈을 수행하고 소프트맥스 레이어를 적용하여 spatial attention map $  &amp;isin; ℝ^{  &amp;times; }$을 계산.&lt;/li&gt;
&lt;li&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;282&quot; data-origin-height=&quot;60&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/8kJ0h/btrdT8LD16b/EZN9VbEKo1TUVEOW1r1Pl0/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/8kJ0h/btrdT8LD16b/EZN9VbEKo1TUVEOW1r1Pl0/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/8kJ0h/btrdT8LD16b/EZN9VbEKo1TUVEOW1r1Pl0/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2F8kJ0h%2FbtrdT8LD16b%2FEZN9VbEKo1TUVEOW1r1Pl0%2Fimg.png&quot; data-origin-width=&quot;282&quot; data-origin-height=&quot;60&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/li&gt;
&lt;li&gt;여기서  ji는 i번째 위치가 j번째 위치에 미치는 영향을 측정. 두 위치의 특성이 더 유사할수록 두 위치간의 상관 관계가 더 커짐.&lt;/li&gt;
&lt;li&gt;새로운 feature 맵 $  &amp;isin; ℝ^{ &amp;times; &amp;times; }$를 생성하기 위해 feature A를 convolution layer 적용하여 $ℝ^{ &amp;times; }$으로 reshape&lt;/li&gt;
&lt;li&gt;그런 다음 D와 S의 전치 사이의 행렬 곱셈을 수행하고 그 결과를 $ℝ^{ &amp;times; &amp;times; }$ 로 reshape.&lt;/li&gt;
&lt;li&gt;마지막으로 스케일 파라미터 &amp;alpha;를 featuresA의 element-wise sum.&lt;/li&gt;
&lt;li&gt;최종 output $E &amp;isin; ℝ^{ &amp;times; &amp;times; }$ 는&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;280&quot; data-origin-height=&quot;56&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/b1QxRH/btrdVjFTVKy/xr4bvF5NKlKCr5fR035zC1/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/b1QxRH/btrdVjFTVKy/xr4bvF5NKlKCr5fR035zC1/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/b1QxRH/btrdVjFTVKy/xr4bvF5NKlKCr5fR035zC1/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2Fb1QxRH%2FbtrdVjFTVKy%2Fxr4bvF5NKlKCr5fR035zC1%2Fimg.png&quot; data-origin-width=&quot;280&quot; data-origin-height=&quot;56&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
와 같이 계산&lt;/li&gt;
&lt;li&gt;alpha는 0으로 초기화 후 사용. 점차 더 많은 가중치를 할당하는 방법을 학습 [28].&lt;/li&gt;
&lt;li&gt;각 위치의 결과 feature E는 모든 위치와 원래 feature에 걸친 feature의 가중치 합&lt;/li&gt;
&lt;li&gt;따라서 global contextual view를 가지고 있으며 spatial attention map에 따라 선택적으로 컨텍스트를 집계.&lt;/li&gt;
&lt;li&gt;similar semantic features은 상호 이익을 달성하여 intra-class compact 및 semantic consistency 향상.&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 data-ke-size=&quot;size20&quot;&gt;&lt;span&gt;&amp;lt; Channel Attention Module&amp;gt;&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;541&quot; data-origin-height=&quot;218&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/bI2wMt/btrdT8SnUm9/zLv3JSU0UTo9FX5W2OXoR0/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/bI2wMt/btrdT8SnUm9/zLv3JSU0UTo9FX5W2OXoR0/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/bI2wMt/btrdT8SnUm9/zLv3JSU0UTo9FX5W2OXoR0/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FbI2wMt%2FbtrdT8SnUm9%2FzLv3JSU0UTo9FX5W2OXoR0%2Fimg.png&quot; data-origin-width=&quot;541&quot; data-origin-height=&quot;218&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;상위 수준의 각 channel map은 class-specific response으로 재평가될 수 있으며 different semantic responses들이 연관&lt;/li&gt;
&lt;li&gt;channel map 간의 interdependencies을 이용하여, 우리는 상호의존적인 feature 맵을 강조하고 specific semantics의 feature 표현을 개선할 수 있음.&lt;/li&gt;
&lt;li&gt;따라서, 채널간 상호의존성을 명시적으로 모델링하기 위해 channel Attention 모듈을 구축&lt;/li&gt;
&lt;li&gt;channel Attention 모듈의 구조는 그림 3(B).&lt;/li&gt;
&lt;li&gt;position Attention 모듈과는 달리 channel Attention map $ &amp;isin;ℝ^{ &amp;times; }$를 원래의 feature $ &amp;isin;ℝ^{ &amp;times; &amp;times; }$로 직접 계산&lt;/li&gt;
&lt;li&gt;특히 A에서 $ℝ^{ &amp;times; }$ 으로 모양을 변경한 후 $A$와$ ^ $사이에 행렬 곱셈을 수행.&lt;/li&gt;
&lt;li&gt;마지막으로, channel Attention map $ &amp;isin;ℝ^{ &amp;times; }$를 얻기 위해 소프트 맥스 레이어를 적용&lt;/li&gt;
&lt;li&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;303&quot; data-origin-height=&quot;65&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/dJqUuB/btrdSdGKDRZ/Km9clzLhfwazJ2hLpRTBd0/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/dJqUuB/btrdSdGKDRZ/Km9clzLhfwazJ2hLpRTBd0/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/dJqUuB/btrdSdGKDRZ/Km9clzLhfwazJ2hLpRTBd0/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FdJqUuB%2FbtrdSdGKDRZ%2FKm9clzLhfwazJ2hLpRTBd0%2Fimg.png&quot; data-origin-width=&quot;303&quot; data-origin-height=&quot;65&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/li&gt;
&lt;li&gt;$x_{ji}$: i번째 채널이 j번째 채널에 미치는 영향을 측정&lt;/li&gt;
&lt;li&gt;X와 $ ^ $ 사이에 행렬 곱셈을 하여 그 결과를 $ℝ^{ &amp;times; &amp;times; }$ 로 reshape, 그 결과에 척도 파라미터 &amp;beta;를 곱하여 A와 element-wise-sum을 수행하여 최종 출력 $  &amp;isin; ℝ^{ &amp;times; &amp;times; }$ 를 구함&lt;/li&gt;
&lt;li&gt;여기서 &amp;beta;는 0부터 점차 가중치를 학습.&lt;/li&gt;
&lt;li&gt;방정식 4는 각 채널의 최종 특징이 모든 채널의 특징과 원래의 특징에 대한 가중치 합이라는 것을 보여주며, feature 맵 사이의 장거리 의미 의존성을 모델링 &amp;gt; feature 차별성을 높이는 데 도움.&lt;/li&gt;
&lt;li&gt;두 채널의 관계를 계산하기 전에 feature을 embed하는 데 Convolution를 사용하지 않음.&lt;/li&gt;
&lt;li&gt;이는 서로 다른 채널 맵 간의 관계를 유지할 수 있기 때문&lt;/li&gt;
&lt;li&gt;또한 global pooling 또는 encoding layer에 의해 채널 관계를 탐색하는 최근 연구[27]와는 달리, 채널 상관 관계를 모델링하기 위해 모든 대응 위치의 공간 정보를 활용&lt;/li&gt;
&lt;/ul&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h4 data-ke-size=&quot;size20&quot;&gt;&lt;span&gt;&amp;lt; Attention Module Embedding with Networks&amp;gt;&lt;/span&gt;&lt;/h4&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;두 가지 Attention 모듈의 feature을 통합.&lt;/li&gt;
&lt;li&gt;특히, 두 개의 Attention 모듈의 출력을 Convolution 계층에 의해 변환하고 요소별 합계를 수행하여 feature fusion을 수행하면 변환 레이어가 따라 최종 예측 맵을 생성.&lt;/li&gt;
&lt;li&gt;더 많은 GPU 메모리가 필요한 계단식 작업을 채택하지 않음.&lt;/li&gt;
&lt;li&gt;Attention 모듈은 단순해서 FCN과 같은 기존 모듈에 직접 삽입 가능.&lt;/li&gt;
&lt;li&gt;매개 변수를 너무 많이 늘리지 않고 feature 표현을 효과적으로 강화.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;&lt;span&gt;Experiments&lt;/span&gt;&lt;/h3&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;평가를 위해 , Cityscapes 데이터 세트[5], PASCAL VOC2012[7], PASCAL Context 데이터 세트 [14] 및 COCO Stuff 데이터 세트에 대한 포괄적인 실험 수행&lt;/li&gt;
&lt;li&gt;실험 결과에 따르면 DANet 은 세 개의 데이터 세트에서 state of the art performance 달성&lt;/li&gt;
&lt;/ul&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h4 data-ke-size=&quot;size20&quot;&gt;&lt;span&gt;&amp;lt;Datasets and Implementation Details&amp;gt;&lt;/span&gt;&lt;/h4&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;Cityscapes : 50 개 도시에서 캡처된 5,000 개의 이미지. 각 이미지는 2048 &amp;times; 1024 픽셀, 19 개 semantic 클래스의 고품질 픽셀 레벨 레이블 존재. training 에는 2,979 개의 영상이 있고 , validation 에는 500 개의 영상이 있으며 , test 세트에는 1,525 개의 영상이 있음. coarse 데이터를 사용하지 않음.&amp;nbsp;&lt;/li&gt;
&lt;li&gt;PASCAL VOC 2012 : training 이미지 10,582 개 , validation 이미지 1,449 개 , test 이미지 1,456 개가 포함 . 여기에는 20 개의 foreground 객체 클래스와 1 개의 background 클래스가 포함&lt;/li&gt;
&lt;li&gt;PASCAL Context : 전체 씬에 대한 자세한 semantic 레이블 제공. training 4,998 개의 이미지와 test 5,105 개의 이미지가 포함 . 가장 빈도가 높은 59 개 클래스에서 하나의 배경 범주 총 60 개 클래스와 함께 평가&lt;/li&gt;
&lt;li&gt;COCO Stuff : training 이미지 9,000 개와 test 이미지 1,000 개가 포함. 각 픽셀에 80 개의 개체와 91 개의 항목에 주석을 달아 171 개의 범주에 대한 결과 보고&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;386&quot; data-origin-height=&quot;232&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/bwFXCD/btrdPJlHGf3/rkh1k9XUpKYK6IvVFq9GK0/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/bwFXCD/btrdPJlHGf3/rkh1k9XUpKYK6IvVFq9GK0/img.png&quot; data-alt=&quot;그림4 : cityscape validation set 에서의 position attention module 적용 결과&amp;amp;amp;nbsp;&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/bwFXCD/btrdPJlHGf3/rkh1k9XUpKYK6IvVFq9GK0/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FbwFXCD%2FbtrdPJlHGf3%2Frkh1k9XUpKYK6IvVFq9GK0%2Fimg.png&quot; data-origin-width=&quot;386&quot; data-origin-height=&quot;232&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;figcaption&gt;그림4 : cityscape validation set 에서의 position attention module 적용 결과&amp;nbsp;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;380&quot; data-origin-height=&quot;219&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/dp6tAo/btrdSRJ0YLh/8NRDY2H3kJijDwOZ6Ig8A1/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/dp6tAo/btrdSRJ0YLh/8NRDY2H3kJijDwOZ6Ig8A1/img.png&quot; data-alt=&quot;그림5 : cityscape validation set 에서의 channel attention module 적용 결과&amp;amp;amp;nbsp;&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/dp6tAo/btrdSRJ0YLh/8NRDY2H3kJijDwOZ6Ig8A1/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2Fdp6tAo%2FbtrdSRJ0YLh%2F8NRDY2H3kJijDwOZ6Ig8A1%2Fimg.png&quot; data-origin-width=&quot;380&quot; data-origin-height=&quot;219&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;figcaption&gt;그림5 : cityscape validation set 에서의 channel attention module 적용 결과&amp;nbsp;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p style=&quot;position: absolute;&quot; data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size18&quot;&gt;&lt;span&gt;Pytorch &lt;/span&gt;&lt;span&gt;기반 구현&lt;/span&gt;&lt;/p&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;[4,27] 에 따라 , 우리는 초기 학습률에학습률에$$(1&amp;minus;\frac{ }{ \_ })^{0.9}$$를곱하는 poly learning rate policy 를 채택&lt;/li&gt;
&lt;li&gt;Cityscapes 데이터셋의 경우 기본 학습률은 0.01 로 설정. Momentum 와 weight decay coefficients 는 각각 0.9 와 0.0001 로 설정 .&lt;/li&gt;
&lt;li&gt;Synchronized BN 을 사용하여 모델을 training. 배치 크기는 Cityscapes 의 경우 8 로 설정되고&lt;/li&gt;
&lt;li&gt;다른 데이터셋의 경우 16 으로 설정&lt;/li&gt;
&lt;li&gt;multi scale augmentation 을 채택할 때 training 시간을 COCO Stuff 의 경우 180Epoch, 기타&lt;/li&gt;
&lt;li&gt;데이터셋의 경우 240Epoch 로 설정 .&lt;/li&gt;
&lt;li&gt;Deeplab에 이어 , 우리는 두 개의 Attention 모듈이 모두 사용될 때 네트워크 끝에서 다중 손실을 채택&lt;/li&gt;
&lt;li&gt;데이터 augmentation 을 위해 Cityscapes 데이터 세트에 대한 ablation study training 중에 무작&lt;/li&gt;
&lt;li&gt;위 cropping ( cropsize 768 ) 와 random left right flipping 을 적용&lt;/li&gt;
&lt;/ul&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h4 data-ke-size=&quot;size20&quot;&gt;&lt;span&gt;&amp;lt; Ablation Study for Attention Modules &amp;gt;&lt;/span&gt;&lt;/h4&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;better scene understanding 의 장거리 의존성을 캡처하기 위해 dilated 네트워크 위에 이중 Attention 모듈을 사용 .&lt;/li&gt;
&lt;li&gt;Attention 모듈의 성능을 검증하기 위해 표 1 의 다양한 설정으로 실험을 수행&lt;/li&gt;
&lt;li&gt;표1 과 같이 Attention 모듈은 성능을 향상시킴 .&lt;/li&gt;
&lt;/ul&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;368&quot; data-origin-height=&quot;252&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/DwxCA/btrdUDR6AkD/KYd2KAbJkQSPpYtK8w55v1/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/DwxCA/btrdUDR6AkD/KYd2KAbJkQSPpYtK8w55v1/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/DwxCA/btrdUDR6AkD/KYd2KAbJkQSPpYtK8w55v1/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FDwxCA%2FbtrdUDR6AkD%2FKYd2KAbJkQSPpYtK8w55v1%2Fimg.png&quot; data-origin-width=&quot;368&quot; data-origin-height=&quot;252&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;기본 FCN(ResNet50) 과 비교했을 때 position Attention 모듈을 채택하면 Mean IoU 에서 75.74% 의 결과를 얻을 수 있어 5.71% 의 개선 효과를 얻을 수 있음 .&lt;/li&gt;
&lt;li&gt;channel contextual module 을 개별적으로 채택하는 것은 기준치를 4. 25% 이상 증가 .&lt;/li&gt;
&lt;li&gt;두 Attention 모듈을 함께 통합하면 성능이 76.34% 로 더욱 향상 .&lt;/li&gt;
&lt;li&gt;또한 , pre-trained ResNet 101 를 채택할 경우 , 두 개의 Attention 모듈을 갖춘 네트워크는 기준 모델에 비해 분할 성능을 5.03% 향상시킵니다 .&lt;/li&gt;
&lt;li&gt;position Attention 모듈의 효과는 그림 4에 시각화 .&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;386&quot; data-origin-height=&quot;232&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/bwFXCD/btrdPJlHGf3/rkh1k9XUpKYK6IvVFq9GK0/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/bwFXCD/btrdPJlHGf3/rkh1k9XUpKYK6IvVFq9GK0/img.png&quot; data-alt=&quot;그림4 : cityscape validation set 에서의 position attention module 적용 결과&amp;amp;amp;nbsp;&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/bwFXCD/btrdPJlHGf3/rkh1k9XUpKYK6IvVFq9GK0/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FbwFXCD%2FbtrdPJlHGf3%2Frkh1k9XUpKYK6IvVFq9GK0%2Fimg.png&quot; data-origin-width=&quot;386&quot; data-origin-height=&quot;232&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;figcaption&gt;그림4 : cityscape validation set 에서의 position attention module 적용 결과&amp;nbsp;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/li&gt;
&lt;li&gt;첫 번째 행의 'pole' 과 두 번째 행의 'sidewalk'와 같은 곳에 position Attention 모듈을 사용하면 일부 세부사항과 객체 경계가 더 명확&lt;/li&gt;
&lt;li&gt;로컬 feature 에 대한 선택적 fusion 은 세부사항의 구별을 강화 .&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;380&quot; data-origin-height=&quot;219&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/dp6tAo/btrdSRJ0YLh/8NRDY2H3kJijDwOZ6Ig8A1/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/dp6tAo/btrdSRJ0YLh/8NRDY2H3kJijDwOZ6Ig8A1/img.png&quot; data-alt=&quot;그림5 : cityscape validation set 에서의 channel attention module 적용 결과&amp;amp;amp;nbsp;&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/dp6tAo/btrdSRJ0YLh/8NRDY2H3kJijDwOZ6Ig8A1/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2Fdp6tAo%2FbtrdSRJ0YLh%2F8NRDY2H3kJijDwOZ6Ig8A1%2Fimg.png&quot; data-origin-width=&quot;380&quot; data-origin-height=&quot;219&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;figcaption&gt;그림5 : cityscape validation set 에서의 channel attention module 적용 결과&amp;nbsp;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;그림 5 는 채널 Attention 모듈을 통해 첫 번째 및 세 번째 행의 bus 와 같이 일부 잘못 분류된 범주가 현재 올바르게 분류되었음을 보여줌 .&lt;/li&gt;
&lt;li&gt;Channel map 간의 선택적 통합은 컨텍스트 정보를 캡처하는 데 도움 .&lt;/li&gt;
&lt;li&gt;일관성이 확실히 개선&lt;/li&gt;
&lt;/ul&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;&amp;gt; Dilated FCN&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;1280&quot; data-origin-height=&quot;373&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/pLWLT/btrdU6fmKSl/9uUkTKXvpIOfZ9tHwrEogK/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/pLWLT/btrdU6fmKSl/9uUkTKXvpIOfZ9tHwrEogK/img.png&quot; data-alt=&quot;출처 :&amp;amp;amp;nbsp;FastFCN: Rethinking Dilated Convolution in the Backbone for Semantic Segmentation&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/pLWLT/btrdU6fmKSl/9uUkTKXvpIOfZ9tHwrEogK/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FpLWLT%2FbtrdU6fmKSl%2F9uUkTKXvpIOfZ9tHwrEogK%2Fimg.png&quot; data-origin-width=&quot;1280&quot; data-origin-height=&quot;373&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;figcaption&gt;출처 :&amp;nbsp;FastFCN: Rethinking Dilated Convolution in the Backbone for Semantic Segmentation&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;파란 &lt;/span&gt;&lt;span&gt;layer : downsampling&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span&gt;주황 &lt;/span&gt;&lt;span&gt;layer : dilated convolutions&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h4 data-ke-size=&quot;size20&quot;&gt;&lt;span&gt;&amp;lt; Ablation Study for Improvement Strategies&amp;gt;&lt;/span&gt;&lt;/h4&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;Deeplabv3[4]에 이어 성능을 더욱 개선하기 위해 동일한 전략을 채택 .&lt;/li&gt;
&lt;li&gt;1) DA( Data augmentation) augmentation): 랜덤 스케일링을 사용한 데이터 augmentation&lt;/li&gt;
&lt;li&gt;2) 다중 그리드 Multi Grid)Grid): 마지막 ResNet 블록에 다양한 크기의 그리드 계층을 적용&lt;/li&gt;
&lt;li&gt;3)MS(Map scaling?): 8 개 이미지 스케일 {0.5, 0.75, 1, 1.25, 1.5, 1.75} 의 분할 확률 맵을 평균화&lt;/li&gt;
&lt;/ul&gt;
&lt;p style=&quot;position: absolute;&quot; data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;368&quot; data-origin-height=&quot;252&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/DwxCA/btrdUDR6AkD/KYd2KAbJkQSPpYtK8w55v1/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/DwxCA/btrdUDR6AkD/KYd2KAbJkQSPpYtK8w55v1/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/DwxCA/btrdUDR6AkD/KYd2KAbJkQSPpYtK8w55v1/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FDwxCA%2FbtrdUDR6AkD%2FKYd2KAbJkQSPpYtK8w55v1%2Fimg.png&quot; data-origin-width=&quot;368&quot; data-origin-height=&quot;252&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;랜덤 dilated 을 통한 데이터 dilated 은 성능을 거의 1.26% 향상 .&lt;/li&gt;
&lt;li&gt;이는 training 데이터의 scale 다양성을 강화함으로써 네트워크 이점을 얻을 수 있음&lt;/li&gt;
&lt;li&gt;사전 훈련된 네트워크의 더 나은 feature 표현을 얻기 위해 멀티그리드를 채택하고 있으며 , 이는 1.11% 의 추가 개선 .&lt;/li&gt;
&lt;li&gt;마지막으로 , Segmentation map fusion 은 성능이 81.50% 로 더욱 향상되어 잘 알려진 방법 Deeplabv3[4] (cityscape val set 에서 79.30%) 을 2.20% 능가&lt;/li&gt;
&lt;/ul&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h4 data-ke-size=&quot;size20&quot;&gt;&lt;span&gt;&amp;lt; Visualization of Attention Module&amp;gt;&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;789&quot; data-origin-height=&quot;296&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/kyO3c/btrdUnIsaPq/KdHOOmBzNt0hVKQqeDYab1/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/kyO3c/btrdUnIsaPq/KdHOOmBzNt0hVKQqeDYab1/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/kyO3c/btrdUnIsaPq/KdHOOmBzNt0hVKQqeDYab1/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FkyO3c%2FbtrdUnIsaPq%2FKdHOOmBzNt0hVKQqeDYab1%2Fimg.png&quot; data-origin-width=&quot;789&quot; data-origin-height=&quot;296&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;position Attention 를 위해 , 전체적인 self Attention map 는 (H &amp;times; W) &amp;times; (H &amp;times; W) 크기로 , 이미지의 각 특정 지점에 대해 (H &amp;times; W) 크기의 해당 하위 Attention map 이 있음을 의미 .&lt;/li&gt;
&lt;li&gt;그림 6 에서는 각 입력 이미지에 대해 두 점 (#1 및 #2 로 표시) 을 선택하고 해당 하위 Attention 맵을 각각 2 열과 3 열에 표시 .&lt;/li&gt;
&lt;li&gt;position Attention 모듈이 명확한 semantic 유사성과 장거리 관계를 포착할 수 있다는 것을 관찰&lt;/li&gt;
&lt;li&gt;예) 첫 번째 행에서 빨간색 포인트 #1 은 건물에 표시되며 Attention map(2 열 는 건물이 있는 대부분의 영역을 강조 . 더욱이, 하위 Attention map 에서, 경계 중 일부는 #1 지점으로부터 멀리 떨어져 있더라도 경계가 매우 명확 .&lt;/li&gt;
&lt;li&gt;예) 포인트 #2 는 Attention map 이 자동차 로 라벨이 지정된 대부분의 위치에 집중 . 두 번째 행에서는 해당 픽셀 수가 적더라도 글로벌 영역의 'traffic 과 ' 을 동일하게 유지 .&lt;/li&gt;
&lt;li&gt;예) 세 번째 행은 'vegetation' 와 &amp;lsquo;person&amp;rsquo; class 를 위한 것 . 특히 포인트 #2 는 가까운 &amp;lsquo;rider&amp;rsquo; class 에는 대응하지 않지만 , &amp;lsquo;person&amp;rsquo; class 에는 대응&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 data-ke-size=&quot;size20&quot;&gt;&lt;span&gt;&amp;lt; Comparing with State of the art&amp;gt;&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;783&quot; data-origin-height=&quot;308&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/cdaKMf/btrdQt3Lbpd/ttpAy8t1qiF5XAhRsakwN0/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/cdaKMf/btrdQt3Lbpd/ttpAy8t1qiF5XAhRsakwN0/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/cdaKMf/btrdQt3Lbpd/ttpAy8t1qiF5XAhRsakwN0/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FcdaKMf%2FbtrdQt3Lbpd%2FttpAy8t1qiF5XAhRsakwN0%2Fimg.png&quot; data-origin-width=&quot;783&quot; data-origin-height=&quot;308&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;Cityscapes 테스트 세트의 기존 방법과 추가로 비교 . 특히 , 주석이 달린 데이터만으로 DANet 101 을 training 후 테스트 결과를 공식 평가 서버에 제출 .&lt;/li&gt;
&lt;li&gt;DANet 은 dominantly advantage. 를 가진 기존 접근 방식을 능가. 특히 , PSANet은 동일한&lt;/li&gt;
&lt;li&gt;백본 ResNet 101을 사용했음에도 성능이 좋음. 더 강력한 사전 훈련된 모델을 사용하는 DenseASPP 까지 능가&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 data-ke-size=&quot;size20&quot;&gt;&lt;span&gt;&amp;lt; Results on PASCAL VOC 2012 Dataset&amp;gt;&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;383&quot; data-origin-height=&quot;172&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/xgitx/btrdPJeZBci/yZHfp8IORz7jhjtRnNH4W0/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/xgitx/btrdPJeZBci/yZHfp8IORz7jhjtRnNH4W0/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/xgitx/btrdPJeZBci/yZHfp8IORz7jhjtRnNH4W0/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2Fxgitx%2FbtrdPJeZBci%2FyZHfp8IORz7jhjtRnNH4W0%2Fimg.png&quot; data-origin-width=&quot;383&quot; data-origin-height=&quot;172&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;추가 효과 평가를 위한 Pascal VOC2012 dataset 에 대한 실험.&lt;/li&gt;
&lt;li&gt;Pascal VOC 의 Quantitative result 는 2012 년 val 세트에서 보여줌 . DANet 50 은 3.3% 를 초과하는 성능향상&lt;/li&gt;
&lt;li&gt;더 깊은 ResNet 101 모델 채택시 mean IoU 80.4% 을 달성&lt;/li&gt;
&lt;li&gt;[4, 27, 29] 에 이어 PASCAL VOC 2012 training 세트에서도 모델을 더 잘 fine tuning. 테스트 세트에 대한 PASCAL VOC 2012 의 결과는 표 5 에 나와있음&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;388&quot; data-origin-height=&quot;249&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/byxJ6C/btrdSc8O2Np/kmQMizhkQR5H8j8rFZyIqK/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/byxJ6C/btrdSc8O2Np/kmQMizhkQR5H8j8rFZyIqK/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/byxJ6C/btrdSc8O2Np/kmQMizhkQR5H8j8rFZyIqK/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FbyxJ6C%2FbtrdSc8O2Np%2FkmQMizhkQR5H8j8rFZyIqK%2Fimg.png&quot; data-origin-width=&quot;388&quot; data-origin-height=&quot;249&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;h4 data-ke-size=&quot;size20&quot;&gt;&lt;span&gt;&amp;lt; Results on PASCAL Context Dataset&amp;gt;&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;376&quot; data-origin-height=&quot;301&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/z3zCt/btrdQCmpX1D/tixmw0ldSwvNHSEtmlmpk1/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/z3zCt/btrdQCmpX1D/tixmw0ldSwvNHSEtmlmpk1/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/z3zCt/btrdQCmpX1D/tixmw0ldSwvNHSEtmlmpk1/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2Fz3zCt%2FbtrdQCmpX1D%2Ftixmw0ldSwvNHSEtmlmpk1%2Fimg.png&quot; data-origin-width=&quot;376&quot; data-origin-height=&quot;301&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;PASCAL Context 에 대한 실험을 수행하여 방법의 효과를 추가로 평가 .&lt;/li&gt;
&lt;li&gt;PASCAL VOC 2012 에 대해 동일한 training 및 test설정을 채택 .&lt;/li&gt;
&lt;li&gt;기준 (dilated FCN 50) 은 평균 IOU 44.3% 를 달성. DANet50 은 성능을 50.1% 로 향상. 깊이 있는 pre-training 네트워크 ResNet101 을 통해 , 우리 모델 결과는 이전 방법들을 큰 차이로 능가하는 Mean IoU 52.6% 를 달성 .&lt;/li&gt;
&lt;li&gt;Deeplab v2 와 RefineNet 은 서로 다른 방식의 변환 또는 다른 단계의 인코더에 의한 멀티스케일 feature fusion 을 채택. 또한 추가 COCO 데이터로 모델을 training하거나 Segmentation 결과를 개선하기 위해 심층 모델을 채택 .&lt;/li&gt;
&lt;li&gt;기존 방식과 달리 , global dependencies을 명시적으로 포착하기 위해 Attention 모듈을 도입하고 , 더 나은 성능 달성&lt;/li&gt;
&lt;/ul&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h4 data-ke-size=&quot;size20&quot;&gt;&lt;span&gt;&amp;lt; Results on COCO Stuff Dataset&amp;gt;&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-origin-width=&quot;373&quot; data-origin-height=&quot;254&quot; data-ke-mobilestyle=&quot;widthOrigin&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/oFWuJ/btrdPBubAYh/elxaXNnXyzxDKtmCA01P51/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/oFWuJ/btrdPBubAYh/elxaXNnXyzxDKtmCA01P51/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/oFWuJ/btrdPBubAYh/elxaXNnXyzxDKtmCA01P51/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FoFWuJ%2FbtrdPBubAYh%2FelxaXNnXyzxDKtmCA01P51%2Fimg.png&quot; data-origin-width=&quot;373&quot; data-origin-height=&quot;254&quot; data-ke-mobilestyle=&quot;widthOrigin&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;제안된 네트워크의 일반화를 검증하기 위해 COCO Stuff 에 대한 실험도 수행 .&lt;/li&gt;
&lt;li&gt;그 결과 , 우리의 모델은 이러한 방법을 큰 차이로 능가하는 Mean IoU 에서 39.7% 를 달성 .&lt;/li&gt;
&lt;li&gt;비교 방법 중 DAG RNN[18]은 2D 이미지용 chain RNN 을 활용하여 풍부한 position dependencies을 모델링&lt;/li&gt;
&lt;li&gt;Ding et al.[6] 눈에 띄지 않는 객체 및 배경 물질 segmentation을개선하기 위해 디코더 단계에서 gating mechanism 채택 .&lt;/li&gt;
&lt;li&gt;우리의 방법은 보다 효과적으로 long-range context information를 포착하고 scene segmentation에서 더 나은 feature representation을 배울 수 있음&lt;/li&gt;
&lt;/ul&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;&lt;span&gt;conclusion&lt;/span&gt;&lt;/h3&gt;
&lt;ul style=&quot;list-style-type: disc;&quot; data-ke-list-type=&quot;disc&quot;&gt;
&lt;li&gt;self Attention mechanism 을 이용하여 local semantic features 을 적응적으로 통합하는 Scene Segmentation 를 위한 DANet(Dual Attention Network) 을 제시 .&lt;/li&gt;
&lt;li&gt;특히 , spatial and channel dimensions 의 global dependencies 을 각각 포착하기 위한 position attention module 과 channel attention module 을 소개 .&lt;/li&gt;
&lt;li&gt;ablation experiments 에서는 dual Attention 모듈이 long range contextual information 를 효과적으로 캡처하고 보다 정밀한 분할 결과를 제공한다는 것을 보여줌 .&lt;/li&gt;
&lt;li&gt;Attention 네트워크는 4 개의 Scene Segmentation 데이터셋(Cityscapes, InPascal VOC 2012, Pascal Context, COCO Stuff) 에서 일관되게 뛰어난 성능을 달성&lt;/li&gt;
&lt;li&gt;추가로 , 컴퓨팅 복잡성을 줄이고 모델의 견고성을 향상시키는 것이 중요하며 , 향후 작업에서 연구 할 것&lt;/li&gt;
&lt;/ul&gt;</description>
      <category>인공지능/Segmentation</category>
      <category>ComputerVision</category>
      <category>DANet</category>
      <category>imageattention</category>
      <category>segmentation</category>
      <category>SemanticSegmentation</category>
      <category>Vision</category>
      <category>visionAttention</category>
      <category>논문리뷰</category>
      <category>어텐션</category>
      <category>이미지어텐션</category>
      <author>통경</author>
      <guid isPermaLink="true">https://stat-cbc.tistory.com/30</guid>
      <comments>https://stat-cbc.tistory.com/30#entry30comment</comments>
      <pubDate>Fri, 3 Sep 2021 12:16:41 +0900</pubDate>
    </item>
    <item>
      <title>Semantic Segmentation 기초 개념</title>
      <link>https://stat-cbc.tistory.com/29</link>
      <description>&lt;p data-ke-size=&quot;size16&quot;&gt;우선, Semantic Segmentation에 대해 알아보기 이전에, Computer Vision의 대표적 Task 2가지인 Object Detection 과 Image Segmentation 의 차이에 대해 알고있어야 합니다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-origin-width=&quot;1094&quot; data-origin-height=&quot;464&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/q39lH/btq93oSCJzy/8wTS0rIAwkikbFBmGPbbyk/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/q39lH/btq93oSCJzy/8wTS0rIAwkikbFBmGPbbyk/img.png&quot; data-alt=&quot;출처 :&amp;amp;nbsp;https://medium.com/zylapp/review-of-deep-learning-algorithms-for-object-detection-c1f3d437b852&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/q39lH/btq93oSCJzy/8wTS0rIAwkikbFBmGPbbyk/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2Fq39lH%2Fbtq93oSCJzy%2F8wTS0rIAwkikbFBmGPbbyk%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;1094&quot; height=&quot;464&quot; data-origin-width=&quot;1094&quot; data-origin-height=&quot;464&quot;/&gt;&lt;/span&gt;&lt;figcaption&gt;출처 :&amp;nbsp;https://medium.com/zylapp/review-of-deep-learning-algorithms-for-object-detection-c1f3d437b852&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;해당 그림을 보면, Object Detection은 여러 객체(Multiple Objects)를 감싸는 Bounding Box(테두리 박스)를 각각 만드는 Localization을 수행하고, 이 Bounding Box가 가지는 객체(class)가 무엇인지에 대해 Classification 을 수행합니다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;반면 Segmentation은 Bounding Box(테두리 박스)없이 객체의 포토샵 누끼를 따듯 경계선을 정확히 분할합니다. 위의 사진에서는 Instance Segmentation으로 언급되어 있지만, 정확히는 Semantic Segmentation입니다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-origin-width=&quot;1280&quot; data-origin-height=&quot;501&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/cf62c0/btq91MNdtFs/UHLhDk03lrDURDfuVXYJ60/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/cf62c0/btq91MNdtFs/UHLhDk03lrDURDfuVXYJ60/img.png&quot; data-alt=&quot;출처 :&amp;amp;nbsp;https://www.jeremyjordan.me/evaluating-image-segmentation-models/&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/cf62c0/btq91MNdtFs/UHLhDk03lrDURDfuVXYJ60/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2Fcf62c0%2Fbtq91MNdtFs%2FUHLhDk03lrDURDfuVXYJ60%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;1280&quot; height=&quot;501&quot; data-origin-width=&quot;1280&quot; data-origin-height=&quot;501&quot;/&gt;&lt;/span&gt;&lt;figcaption&gt;출처 :&amp;nbsp;https://www.jeremyjordan.me/evaluating-image-segmentation-models/&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;Instance Segmentation은 위의 우측 사진처럼 단순히 객체만 분류하는 것이 아니라, 같은 객체여도 서로 다른 instance 를 분류해주는 점이 Semantic Segmentation과 차이가 있습니다. 그렇기 때문에 Instance segmentation과 Semantic Segmentation은 명확하게는 다른 용어입니다. 추후에는 Semantic Segmentation만 다룰 예정이므로 이 부분에 대해 이해가 안되셨어도 넘어가시면 됩니다.&amp;nbsp;&lt;/p&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;Semantic Segmantation이란?&amp;nbsp;&lt;/h3&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-origin-width=&quot;500&quot; data-origin-height=&quot;250&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/bIZwx3/btq97PBGI7I/9dI4kEKZd6SdRBrSnokepk/img.gif&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/bIZwx3/btq97PBGI7I/9dI4kEKZd6SdRBrSnokepk/img.gif&quot; data-alt=&quot;출처 :&amp;amp;nbsp;https://towardsdatascience.com/semantic-segmentation-with-deep-learning-a-guide-and-code-e52fc8958823&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/bIZwx3/btq97PBGI7I/9dI4kEKZd6SdRBrSnokepk/img.gif&quot; srcset=&quot;https://blog.kakaocdn.net/dn/bIZwx3/btq97PBGI7I/9dI4kEKZd6SdRBrSnokepk/img.gif&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;500&quot; height=&quot;250&quot; data-origin-width=&quot;500&quot; data-origin-height=&quot;250&quot;/&gt;&lt;/span&gt;&lt;figcaption&gt;출처 :&amp;nbsp;https://towardsdatascience.com/semantic-segmentation-with-deep-learning-a-guide-and-code-e52fc8958823&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;딥러닝 기반 Semantic Segmentation 은 기존 라이다나 센서 기반이었던 자율주행의 판도를 바꿀 정도로 빠르게 발전하고 있습니다. 자율주행 뿐만아니라 사진 촬영시 사물제외 배경 블러 처리, Crack 탐지를 통한 노후도 측정, 세포 구분(U-Net) 등과 같은 다양한 Task에서 사용됩니다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;Semantic Segmentation은 해당 영상(Image)의 모든 픽셀 각각의 class를 Classification하는 Task입니다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-origin-width=&quot;1046&quot; data-origin-height=&quot;368&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/bOGkoE/btq96EAEzgO/dYXBkZaQv5uC5wFzAWfNJk/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/bOGkoE/btq96EAEzgO/dYXBkZaQv5uC5wFzAWfNJk/img.png&quot; data-alt=&quot;출처 :&amp;amp;nbsp;https://www.jeremyjordan.me/semantic-segmentation/#dilated_convolutions&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/bOGkoE/btq96EAEzgO/dYXBkZaQv5uC5wFzAWfNJk/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FbOGkoE%2Fbtq96EAEzgO%2FdYXBkZaQv5uC5wFzAWfNJk%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;1046&quot; height=&quot;368&quot; data-origin-width=&quot;1046&quot; data-origin-height=&quot;368&quot;/&gt;&lt;/span&gt;&lt;figcaption&gt;출처 :&amp;nbsp;https://www.jeremyjordan.me/semantic-segmentation/#dilated_convolutions&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;Input Image(RGB-색상 3차원(3chanel) 또는 흑백(1 chanel)를 받아 처리를 하는 과정은 논문별로 다르겠지만, Output이 Pixel 별로 Class Label을 가지는 것은 같습니다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-origin-width=&quot;1021&quot; data-origin-height=&quot;294&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/bZrKpj/btq92xWGVG5/V2EtM9DjKxkoWjmNuzF9ik/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/bZrKpj/btq92xWGVG5/V2EtM9DjKxkoWjmNuzF9ik/img.png&quot; data-alt=&quot;출처 :&amp;amp;nbsp;https://www.jeremyjordan.me/semantic-segmentation/#dilated_convolutions&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/bZrKpj/btq92xWGVG5/V2EtM9DjKxkoWjmNuzF9ik/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FbZrKpj%2Fbtq92xWGVG5%2FV2EtM9DjKxkoWjmNuzF9ik%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;1021&quot; height=&quot;294&quot; data-origin-width=&quot;1021&quot; data-origin-height=&quot;294&quot;/&gt;&lt;/span&gt;&lt;figcaption&gt;출처 :&amp;nbsp;https://www.jeremyjordan.me/semantic-segmentation/#dilated_convolutions&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;다만, 모델을 학습시키기 위해서는 Ground Truth라는 픽셀별로 라벨링된 &lt;b&gt;정답&lt;/b&gt; 라벨이 있어야 합니다. Custom data에 라벨링을 하는 경우 다양한 툴을 사용해야 합니다. 가장 많이 알고 계시는 포토샵을 사용하기도 하고 따로 Segmentation Tool을 사용하기도 합니다.&amp;nbsp;&lt;/p&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;Semantic Segmentaion 평가 요소&lt;/h3&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;Semantic Segmentation의 평가요소는 크게 두 가지로 구분해볼 수 있습니다. 바로 성능 측면과, 속도 측면입니다. 자율주행과 같은 실시간 적용이 필요한 Task의 경우엔 아무리 성능이 좋아도 시간이 너무 오래걸리게 되면 사용이 불가하기 때문입니다. (+자율주행과 같은 실시간(real-time) Task에 사용하려면 Fps 기준 최소 30 이상은 나와야 한다고 합니다)&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;1) 성능&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-origin-width=&quot;872&quot; data-origin-height=&quot;333&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/LpPtj/btq99mFN6HF/SrsHVwsNhmz1dIsKdsW3yk/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/LpPtj/btq99mFN6HF/SrsHVwsNhmz1dIsKdsW3yk/img.png&quot; data-alt=&quot;사진 출처 :&amp;amp;nbsp;https://www.jeremyjordan.me/evaluating-image-segmentation-models/&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/LpPtj/btq99mFN6HF/SrsHVwsNhmz1dIsKdsW3yk/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FLpPtj%2Fbtq99mFN6HF%2FSrsHVwsNhmz1dIsKdsW3yk%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;872&quot; height=&quot;333&quot; data-origin-width=&quot;872&quot; data-origin-height=&quot;333&quot;/&gt;&lt;/span&gt;&lt;figcaption&gt;사진 출처 :&amp;nbsp;https://www.jeremyjordan.me/evaluating-image-segmentation-models/&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;그럼, 사람에 대한 Class만 예측한다고 가정할 때 정답 라벨이 왼쪽과 같고, 모델이 예측한 라벨이 오른쪽과 같은 경우에 성능에 관련한 지표를 어떻게 평가하는지 확인해보면, 평소 많이 사용하는 confusion matrix로 생각해보면 다음과 같이 구분할 수 있다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-origin-width=&quot;849&quot; data-origin-height=&quot;374&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/sAs4D/btq97yzXtIH/3FkK5EMBNXQE1s5kIILXf1/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/sAs4D/btq97yzXtIH/3FkK5EMBNXQE1s5kIILXf1/img.png&quot; data-alt=&quot;사진 출처 :&amp;amp;nbsp;https://www.jeremyjordan.me/evaluating-image-segmentation-models/&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/sAs4D/btq97yzXtIH/3FkK5EMBNXQE1s5kIILXf1/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FsAs4D%2Fbtq97yzXtIH%2F3FkK5EMBNXQE1s5kIILXf1%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;849&quot; height=&quot;374&quot; data-origin-width=&quot;849&quot; data-origin-height=&quot;374&quot;/&gt;&lt;/span&gt;&lt;figcaption&gt;사진 출처 :&amp;nbsp;https://www.jeremyjordan.me/evaluating-image-segmentation-models/&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-origin-width=&quot;512&quot; data-origin-height=&quot;130&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/cMoEaq/btq92yBeOqa/mUu476IDBZznIVcI1k4RwK/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/cMoEaq/btq92yBeOqa/mUu476IDBZznIVcI1k4RwK/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/cMoEaq/btq92yBeOqa/mUu476IDBZznIVcI1k4RwK/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FcMoEaq%2Fbtq92yBeOqa%2FmUu476IDBZznIVcI1k4RwK%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;512&quot; height=&quot;130&quot; data-origin-width=&quot;512&quot; data-origin-height=&quot;130&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;모델이 예측한 부분이 Class에 해당하는 부분은 TP(True Positive),&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;모델이 예측하지 않은 부분 중 Class에 해당하지 않는 부분은 TN(True Negative),&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;모델이 예측한 부분이 Class에 해당하지 않는 부분은 FP(False Positive),&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;모델이 예측하지 않은 부분이 Class에 해당하는 부분은 FN(False Negative) 이다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;- IOU(&lt;span style=&quot;color: #313b3f;&quot;&gt;Intersection over Union)&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;IOU를 쉽게 설명하자면&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-origin-width=&quot;392&quot; data-origin-height=&quot;107&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/sYUuZ/btq99mTk9LP/nhZI8JqAXggf1TvqZDIel0/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/sYUuZ/btq99mTk9LP/nhZI8JqAXggf1TvqZDIel0/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/sYUuZ/btq99mTk9LP/nhZI8JqAXggf1TvqZDIel0/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FsYUuZ%2Fbtq99mTk9LP%2FnhZI8JqAXggf1TvqZDIel0%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;392&quot; height=&quot;107&quot; data-origin-width=&quot;392&quot; data-origin-height=&quot;107&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;이다. 즉, 분자는 모델이 정답을 맞춘 부분이고, 분모는 정답 + 모델이 예측한 부분(잘못 예측한 부분도 함께) 이다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;이를 confusion matrix 요소로도 표현해보면&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-origin-width=&quot;325&quot; data-origin-height=&quot;126&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/KBfJx/btq91N6rSvT/TB4wzuw4ChZ1DYSwmhTTZK/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/KBfJx/btq91N6rSvT/TB4wzuw4ChZ1DYSwmhTTZK/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/KBfJx/btq91N6rSvT/TB4wzuw4ChZ1DYSwmhTTZK/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FKBfJx%2Fbtq91N6rSvT%2FTB4wzuw4ChZ1DYSwmhTTZK%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;325&quot; height=&quot;126&quot; data-origin-width=&quot;325&quot; data-origin-height=&quot;126&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;와 같다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;- Pixel Accuracy&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;Pixel Accuracy 는 말 그대로 전체 정답 라벨 중에 모델이 정답 Class 를 맞춘 부분이다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-origin-width=&quot;505&quot; data-origin-height=&quot;155&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/bs8to4/btq99nEHK52/TuHTSTOkxERvyAX6VKlDkK/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/bs8to4/btq99nEHK52/TuHTSTOkxERvyAX6VKlDkK/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/bs8to4/btq99nEHK52/TuHTSTOkxERvyAX6VKlDkK/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2Fbs8to4%2Fbtq99nEHK52%2FTuHTSTOkxERvyAX6VKlDkK%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;505&quot; height=&quot;155&quot; data-origin-width=&quot;505&quot; data-origin-height=&quot;155&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;기본적인 개념은 같으나 이를 Class별, 픽셀별 등으로 나누어서 사용하기도 한다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-origin-width=&quot;820&quot; data-origin-height=&quot;424&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/8ekdg/btq98uc0eRQ/HOWKxLhENm7JKy6DHHFvJ1/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/8ekdg/btq98uc0eRQ/HOWKxLhENm7JKy6DHHFvJ1/img.png&quot; data-alt=&quot;출처 : https://www.jeremyjordan.me/evaluating-image-segmentation-models/&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/8ekdg/btq98uc0eRQ/HOWKxLhENm7JKy6DHHFvJ1/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2F8ekdg%2Fbtq98uc0eRQ%2FHOWKxLhENm7JKy6DHHFvJ1%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;820&quot; height=&quot;424&quot; data-origin-width=&quot;820&quot; data-origin-height=&quot;424&quot;/&gt;&lt;/span&gt;&lt;figcaption&gt;출처 : https://www.jeremyjordan.me/evaluating-image-segmentation-models/&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;이외에도 f1-score 등도 사용하기도 한다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;2) 속도&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;- Fps(&lt;b&gt;F&lt;/b&gt;&lt;span style=&quot;color: #4d5156;&quot;&gt;rames &lt;b&gt;P&lt;/b&gt;er &lt;b&gt;S&lt;/b&gt;econd)&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;color: #4d5156;&quot;&gt;1초당 몇 프레임을 처리할 수 있는지를 말한다. 실시간(real-time) 사용 가능하려면 최소 30 fps 의 성능이 나와야 한다.&amp;nbsp;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;- GFLOPs(&lt;b&gt;G&lt;/b&gt;iga &lt;b&gt;FL&lt;/b&gt;&lt;span style=&quot;color: #202122;&quot;&gt;oating point&lt;span&gt;&amp;nbsp;&lt;/span&gt;&lt;/span&gt;&lt;b&gt;O&lt;/b&gt;&lt;span style=&quot;color: #202122;&quot;&gt;perations&lt;span&gt;&amp;nbsp;&lt;/span&gt;&lt;/span&gt;&lt;b&gt;P&lt;/b&gt;&lt;span style=&quot;color: #202122;&quot;&gt;er&lt;span&gt;&amp;nbsp;&lt;/span&gt;&lt;/span&gt;&lt;b&gt;S&lt;/b&gt;&lt;span style=&quot;color: #202122;&quot;&gt;econd)&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;color: #202122;&quot;&gt;초당 부동소수점 연산이란 의미로 1초동안 수행할 수 있는 부동소수점 연산의 횟수를 기준으로 삼음.&amp;nbsp;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;color: #202122;&quot;&gt;자세한 내용은 위키 페이지(&lt;a href=&quot;https://ko.wikipedia.org/wiki/%ED%94%8C%EB%A1%AD%EC%8A%A4&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot;&gt;https://ko.wikipedia.org/wiki/%ED%94%8C%EB%A1%AD%EC%8A%A4&lt;/a&gt;) 참조&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;- Time 단위(&lt;span style=&quot;color: #373a3c;&quot;&gt;ms etc.)&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;color: #373a3c;&quot;&gt;밀리세컨드(millisecond)를 기준으로 많이 사용함. 프로그램 수행 시간을 측정할 때 초를 대신하는 시간의 단위로 활용한다.&amp;nbsp;&lt;/span&gt;&lt;/p&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;Semantic Segmentaion 구성요소&amp;nbsp;&lt;/h3&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-origin-width=&quot;1021&quot; data-origin-height=&quot;294&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/bZrKpj/btq92xWGVG5/V2EtM9DjKxkoWjmNuzF9ik/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/bZrKpj/btq92xWGVG5/V2EtM9DjKxkoWjmNuzF9ik/img.png&quot; data-alt=&quot;출처 :&amp;amp;nbsp;https://www.jeremyjordan.me/semantic-segmentation/#dilated_convolutions&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/bZrKpj/btq92xWGVG5/V2EtM9DjKxkoWjmNuzF9ik/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FbZrKpj%2Fbtq92xWGVG5%2FV2EtM9DjKxkoWjmNuzF9ik%2Fimg.png&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;1021&quot; height=&quot;294&quot; data-origin-width=&quot;1021&quot; data-origin-height=&quot;294&quot;/&gt;&lt;/span&gt;&lt;figcaption&gt;출처 :&amp;nbsp;https://www.jeremyjordan.me/semantic-segmentation/#dilated_convolutions&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;Segmentation 은 기존 Object Detection 의 VGG, ResNet 등 처럼 수많은 Layer를 사용한다고 성능이 좋게 나오지 않습니다. 왜냐하면, 여러 Layer를 거치게 될 수록 Max, Average Pooling이나 특히 FC(Fully Connected) 과 같은 변형이 발생하게 되면 픽셀의 위치 정보가 손실되기 때문이다. 그렇기 때문에 Segmentation은 FC(Fully Connected) 사용을 지양한다. 대부분의 Sota Segmentation 모델에도 FC(Fully Connected) 레이어가 잘 포함되어 있지 않다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;공간정보를 위해 Max, Average Pooling이나 특히 FC(Fully Connected) 레이어를 모두 삭제하고 Stride 1과 같은 Convolution Layer 만을 사용할 수도 있지만, Parameter 수가 너무 많아 효율이 떨어진다. 이렇게 구성한 모델은 애초에 Fps 가 30은 커녕 0에 수렴할 것이다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;Segmentation에서는 &lt;span style=&quot;color: #292929;&quot;&gt;Downsampling &amp;amp; Upsampling&lt;span&gt; 을 사용한다. Segnet논문 이후에는 Encoder, Decoder라는 이름으로도 많이 사용하는데, &lt;span style=&quot;color: #292929;&quot;&gt;Encoder는 주로 &lt;span style=&quot;color: #292929;&quot;&gt;Downsampling을 여러 번 하는 구조를 담고 있고 &lt;span style=&quot;color: #292929;&quot;&gt;Decoder는 &lt;span style=&quot;color: #292929;&quot;&gt;Upsampling&lt;/span&gt;&lt;span style=&quot;color: #292929;&quot;&gt;&lt;span&gt; 을 여러 번 수행하는 구조를 담고 있다.&amp;nbsp;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style=&quot;color: #292929;&quot;&gt;&lt;span&gt;&lt;span style=&quot;color: #292929;&quot;&gt;&lt;span style=&quot;color: #292929;&quot;&gt;&lt;span style=&quot;color: #292929;&quot;&gt;&lt;span style=&quot;color: #292929;&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;color: #292929;&quot;&gt;- Downsampling 이란, 이미지 Input 의 차원을 줄여 적은 Memory만을 사용하여 모델을 적용할 수 있도록 하는 작업이다. &lt;span style=&quot;color: #292929;&quot;&gt;Convolution 사용시에는 &lt;/span&gt;Stride를 최소 2이상으로 설정해 사용한다.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-origin-width=&quot;395&quot; data-origin-height=&quot;449&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/xLaVN/btq97P9w0hL/ozd5kNkh3Qd5Htz6TVSpj1/img.gif&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/xLaVN/btq97P9w0hL/ozd5kNkh3Qd5Htz6TVSpj1/img.gif&quot; data-alt=&quot;출처 :&amp;amp;nbsp;https://towardsdatascience.com/types-of-convolutions-in-deep-learning-717013397f4d&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/xLaVN/btq97P9w0hL/ozd5kNkh3Qd5Htz6TVSpj1/img.gif&quot; srcset=&quot;https://blog.kakaocdn.net/dn/xLaVN/btq97P9w0hL/ozd5kNkh3Qd5Htz6TVSpj1/img.gif&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;395&quot; height=&quot;449&quot; data-origin-width=&quot;395&quot; data-origin-height=&quot;449&quot;/&gt;&lt;/span&gt;&lt;figcaption&gt;출처 :&amp;nbsp;https://towardsdatascience.com/types-of-convolutions-in-deep-learning-717013397f4d&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;color: #292929;&quot;&gt;하지만 이에 대한 문제도 발생하여 Atrous Convolution(&lt;span style=&quot;color: #666666;&quot;&gt;dilated convolution)&lt;/span&gt;같이 중간중간 건너 뛰는 &lt;span style=&quot;color: #292929;&quot;&gt;Convolution&lt;span&gt; 을 사용하기도 한다.&amp;nbsp;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-ke-mobileStyle=&quot;widthOrigin&quot; data-origin-width=&quot;395&quot; data-origin-height=&quot;381&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/vKNE0/btq98vbYsGo/KZYO33LivimvXsq2zOIunk/img.gif&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/vKNE0/btq98vbYsGo/KZYO33LivimvXsq2zOIunk/img.gif&quot; data-alt=&quot;출처 :&amp;amp;nbsp;https://towardsdatascience.com/types-of-convolutions-in-deep-learning-717013397f4d&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/vKNE0/btq98vbYsGo/KZYO33LivimvXsq2zOIunk/img.gif&quot; srcset=&quot;https://blog.kakaocdn.net/dn/vKNE0/btq98vbYsGo/KZYO33LivimvXsq2zOIunk/img.gif&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot; loading=&quot;lazy&quot; width=&quot;395&quot; height=&quot;381&quot; data-origin-width=&quot;395&quot; data-origin-height=&quot;381&quot;/&gt;&lt;/span&gt;&lt;figcaption&gt;출처 :&amp;nbsp;https://towardsdatascience.com/types-of-convolutions-in-deep-learning-717013397f4d&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;이외에도 다양한 &lt;span style=&quot;color: #292929;&quot;&gt;Downsampling&lt;span&gt; 방법을 사용한다.&amp;nbsp;&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;- &lt;span style=&quot;color: #292929;&quot;&gt;Upsampling&lt;/span&gt;&lt;span style=&quot;color: #292929;&quot;&gt;&lt;span&gt; 이란, &lt;span style=&quot;color: #292929;&quot;&gt;Downsampling&lt;/span&gt;&lt;span style=&quot;color: #292929;&quot;&gt;&lt;span&gt; 에서 축소한만큼 다시 확대를 해서 원본 Input과 같은 차원으로 만들어주는 과정을 말한다.&amp;nbsp;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;color: #292929;&quot;&gt;Upsampling&lt;/span&gt;&lt;span style=&quot;color: #292929;&quot;&gt;&lt;span&gt;&lt;span&gt; 에서도 다양한 방법을 사용하는데, 단순 &lt;span style=&quot;color: #292929;&quot;&gt;Transpose Convolution, &lt;/span&gt;&lt;span style=&quot;color: #292929;&quot;&gt;Strided Transpose Convolution 이외에도 다양한 방법을 사용한다.&amp;nbsp;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;color: #292929;&quot;&gt;&lt;span&gt;&lt;span&gt;&lt;span style=&quot;color: #292929;&quot;&gt;이로써 기본적인 Segmentation 개념에 대해 알아보았고, 이제 추후 다룰 논문들을 통해 어떤 식으로 Segmentation 논문이 발전하고 있는지에 대하여 살펴볼 예정입니다!&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;color: #292929;&quot;&gt;&lt;span&gt;&lt;span&gt;&lt;span style=&quot;color: #292929;&quot;&gt;1) 기본 Semantic Segmentation&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;color: #292929;&quot;&gt;&lt;span&gt;&lt;span&gt;&lt;span style=&quot;color: #292929;&quot;&gt;2) 실시간(Real time) Task를 위한 경량화 &lt;span style=&quot;color: #292929;&quot;&gt;Semantic Segmentation&amp;nbsp;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;color: #292929;&quot;&gt;&lt;span&gt;&lt;span&gt;&lt;span style=&quot;color: #292929;&quot;&gt;&lt;span style=&quot;color: #292929;&quot;&gt;3) Attention계열 &lt;span style=&quot;color: #292929;&quot;&gt;Semantic Segmentation&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;등과 같이 나눠볼 수 있겠습니다.&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;reference : &lt;a href=&quot;https://towardsdatascience.com/semantic-segmentation-with-deep-learning-a-guide-and-code-e52fc8958823&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot;&gt;https://towardsdatascience.com/semantic-segmentation-with-deep-learning-a-guide-and-code-e52fc8958823&lt;/a&gt;&lt;/p&gt;
&lt;figure id=&quot;og_1626790572665&quot; contenteditable=&quot;false&quot; data-ke-type=&quot;opengraph&quot; data-ke-align=&quot;alignCenter&quot; data-og-type=&quot;article&quot; data-og-title=&quot;Semantic Segmentation with Deep Learning&quot; data-og-description=&quot;Most people in the deep learning and computer vision communities understand what image classification is: we want our model to tell us what single object or scene is present in the image&amp;hellip;&quot; data-og-host=&quot;towardsdatascience.com&quot; data-og-source-url=&quot;https://towardsdatascience.com/semantic-segmentation-with-deep-learning-a-guide-and-code-e52fc8958823&quot; data-og-url=&quot;https://towardsdatascience.com/semantic-segmentation-with-deep-learning-a-guide-and-code-e52fc8958823&quot; data-og-image=&quot;https://scrap.kakaocdn.net/dn/CGZBQ/hyKWIAbKJ8/KwZbDL14svHxSLV6K9kYm1/img.jpg?width=500&amp;amp;height=250&amp;amp;face=0_0_500_250,https://scrap.kakaocdn.net/dn/s51zh/hyKWU1EJvy/CXUndKOD9znNQ9ONdiUtP0/img.png?width=60&amp;amp;height=39&amp;amp;face=0_0_60_39,https://scrap.kakaocdn.net/dn/cdjN7U/hyKXNs6hlx/TyWWiaHHZ7leGY7Y3sftr1/img.png?width=60&amp;amp;height=43&amp;amp;face=0_0_60_43&quot;&gt;&lt;a href=&quot;https://towardsdatascience.com/semantic-segmentation-with-deep-learning-a-guide-and-code-e52fc8958823&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; data-source-url=&quot;https://towardsdatascience.com/semantic-segmentation-with-deep-learning-a-guide-and-code-e52fc8958823&quot;&gt;
&lt;div class=&quot;og-image&quot; style=&quot;background-image: url('https://scrap.kakaocdn.net/dn/CGZBQ/hyKWIAbKJ8/KwZbDL14svHxSLV6K9kYm1/img.jpg?width=500&amp;amp;height=250&amp;amp;face=0_0_500_250,https://scrap.kakaocdn.net/dn/s51zh/hyKWU1EJvy/CXUndKOD9znNQ9ONdiUtP0/img.png?width=60&amp;amp;height=39&amp;amp;face=0_0_60_39,https://scrap.kakaocdn.net/dn/cdjN7U/hyKXNs6hlx/TyWWiaHHZ7leGY7Y3sftr1/img.png?width=60&amp;amp;height=43&amp;amp;face=0_0_60_43');&quot;&gt;&amp;nbsp;&lt;/div&gt;
&lt;div class=&quot;og-text&quot;&gt;
&lt;p class=&quot;og-title&quot; data-ke-size=&quot;size16&quot;&gt;Semantic Segmentation with Deep Learning&lt;/p&gt;
&lt;p class=&quot;og-desc&quot; data-ke-size=&quot;size16&quot;&gt;Most people in the deep learning and computer vision communities understand what image classification is: we want our model to tell us what single object or scene is present in the image&amp;hellip;&lt;/p&gt;
&lt;p class=&quot;og-host&quot; data-ke-size=&quot;size16&quot;&gt;towardsdatascience.com&lt;/p&gt;
&lt;/div&gt;
&lt;/a&gt;&lt;/figure&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;color: #777777;&quot;&gt;&lt;a href=&quot;https://www.jeremyjordan.me/evaluating-image-segmentation-models/&quot;&gt;https://www.jeremyjordan.me/evaluating-image-segmentation-models/&lt;/a&gt;&lt;/span&gt;&lt;/p&gt;
&lt;figure id=&quot;og_1626790582005&quot; contenteditable=&quot;false&quot; data-ke-type=&quot;opengraph&quot; data-ke-align=&quot;alignCenter&quot; data-og-type=&quot;article&quot; data-og-title=&quot;Evaluating image segmentation models.&quot; data-og-description=&quot;When evaluating a standard machine learning model, we usually classify our predictions into four categories: true positives, false positives, true negatives, and false negatives. However, for the dense prediction task of image segmentation, it's not immedi&quot; data-og-host=&quot;www.jeremyjordan.me&quot; data-og-source-url=&quot;https://www.jeremyjordan.me/evaluating-image-segmentation-models/&quot; data-og-url=&quot;https://www.jeremyjordan.me/evaluating-image-segmentation-models/&quot; data-og-image=&quot;&quot;&gt;&lt;a href=&quot;https://www.jeremyjordan.me/evaluating-image-segmentation-models/&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; data-source-url=&quot;https://www.jeremyjordan.me/evaluating-image-segmentation-models/&quot;&gt;
&lt;div class=&quot;og-image&quot; style=&quot;background-image: url();&quot;&gt;&amp;nbsp;&lt;/div&gt;
&lt;div class=&quot;og-text&quot;&gt;
&lt;p class=&quot;og-title&quot; data-ke-size=&quot;size16&quot;&gt;Evaluating image segmentation models.&lt;/p&gt;
&lt;p class=&quot;og-desc&quot; data-ke-size=&quot;size16&quot;&gt;When evaluating a standard machine learning model, we usually classify our predictions into four categories: true positives, false positives, true negatives, and false negatives. However, for the dense prediction task of image segmentation, it's not immedi&lt;/p&gt;
&lt;p class=&quot;og-host&quot; data-ke-size=&quot;size16&quot;&gt;www.jeremyjordan.me&lt;/p&gt;
&lt;/div&gt;
&lt;/a&gt;&lt;/figure&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;span style=&quot;color: #777777;&quot;&gt;&lt;a href=&quot;https://medium.com/zylapp/review-of-deep-learning-algorithms-for-object-detection-c1f3d437b852&quot;&gt;https://medium.com/zylapp/review-of-deep-learning-algorithms-for-object-detection-c1f3d437b852&lt;/a&gt;&lt;/span&gt;&lt;/p&gt;
&lt;figure id=&quot;og_1626791211116&quot; contenteditable=&quot;false&quot; data-ke-type=&quot;opengraph&quot; data-ke-align=&quot;alignCenter&quot; data-og-type=&quot;article&quot; data-og-title=&quot;Review of Deep Learning Algorithms for Object Detection&quot; data-og-description=&quot;Why object detection instead of image classification?&quot; data-og-host=&quot;medium.com&quot; data-og-source-url=&quot;https://medium.com/zylapp/review-of-deep-learning-algorithms-for-object-detection-c1f3d437b852&quot; data-og-url=&quot;https://medium.com/zylapp/review-of-deep-learning-algorithms-for-object-detection-c1f3d437b852&quot; data-og-image=&quot;https://scrap.kakaocdn.net/dn/zqRlm/hyKWJZ9cmK/x6MoyQezk3uqKnBMYVMLH0/img.jpg?width=1200&amp;amp;height=499&amp;amp;face=0_0_1200_499,https://scrap.kakaocdn.net/dn/ctuzU4/hyKXNs6zK4/Q9JfMpPMpJSApcdp9KmQpk/img.png?width=60&amp;amp;height=36&amp;amp;face=0_0_60_36,https://scrap.kakaocdn.net/dn/mG4KL/hyKXUyZnau/HB51KYGYgbUkCaRKudPBk1/img.png?width=60&amp;amp;height=34&amp;amp;face=0_0_60_34&quot;&gt;&lt;a href=&quot;https://medium.com/zylapp/review-of-deep-learning-algorithms-for-object-detection-c1f3d437b852&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; data-source-url=&quot;https://medium.com/zylapp/review-of-deep-learning-algorithms-for-object-detection-c1f3d437b852&quot;&gt;
&lt;div class=&quot;og-image&quot; style=&quot;background-image: url('https://scrap.kakaocdn.net/dn/zqRlm/hyKWJZ9cmK/x6MoyQezk3uqKnBMYVMLH0/img.jpg?width=1200&amp;amp;height=499&amp;amp;face=0_0_1200_499,https://scrap.kakaocdn.net/dn/ctuzU4/hyKXNs6zK4/Q9JfMpPMpJSApcdp9KmQpk/img.png?width=60&amp;amp;height=36&amp;amp;face=0_0_60_36,https://scrap.kakaocdn.net/dn/mG4KL/hyKXUyZnau/HB51KYGYgbUkCaRKudPBk1/img.png?width=60&amp;amp;height=34&amp;amp;face=0_0_60_34');&quot;&gt;&amp;nbsp;&lt;/div&gt;
&lt;div class=&quot;og-text&quot;&gt;
&lt;p class=&quot;og-title&quot; data-ke-size=&quot;size16&quot;&gt;Review of Deep Learning Algorithms for Object Detection&lt;/p&gt;
&lt;p class=&quot;og-desc&quot; data-ke-size=&quot;size16&quot;&gt;Why object detection instead of image classification?&lt;/p&gt;
&lt;p class=&quot;og-host&quot; data-ke-size=&quot;size16&quot;&gt;medium.com&lt;/p&gt;
&lt;/div&gt;
&lt;/a&gt;&lt;/figure&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&lt;a href=&quot;https://www.jeremyjordan.me/semantic-segmentation/#dilated_convolutions&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot;&gt;https://www.jeremyjordan.me/semantic-segmentation/#dilated_convolutions&lt;/a&gt;&lt;/p&gt;
&lt;figure id=&quot;og_1626791215866&quot; contenteditable=&quot;false&quot; data-ke-type=&quot;opengraph&quot; data-ke-align=&quot;alignCenter&quot; data-og-type=&quot;article&quot; data-og-title=&quot;An overview of semantic image segmentation.&quot; data-og-description=&quot;In this post, I'll discuss how to use convolutional neural networks for the task of semantic image segmentation. Image segmentation is a computer vision task in which we label specific regions of an image according to what's being shown. &amp;quot;What's in this im&quot; data-og-host=&quot;www.jeremyjordan.me&quot; data-og-source-url=&quot;https://www.jeremyjordan.me/semantic-segmentation/#dilated_convolutions&quot; data-og-url=&quot;https://www.jeremyjordan.me/semantic-segmentation/&quot; data-og-image=&quot;https://scrap.kakaocdn.net/dn/cOsUsU/hyKWSW5JD6/CM2DYhWciDNFZkbCTZjj6k/img.png?width=2448&amp;amp;height=1780&amp;amp;face=0_0_2448_1780,https://scrap.kakaocdn.net/dn/BzvMg/hyKXSBboye/Gw0goJMcKmrBEeKm7tH0qk/img.png?width=2448&amp;amp;height=1780&amp;amp;face=0_0_2448_1780,https://scrap.kakaocdn.net/dn/erhIww/hyKWNO1CdY/wp8kWr4t4Xq9P7sMI6UA61/img.png?width=1194&amp;amp;height=693&amp;amp;face=0_0_1194_693&quot;&gt;&lt;a href=&quot;https://www.jeremyjordan.me/semantic-segmentation/#dilated_convolutions&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; data-source-url=&quot;https://www.jeremyjordan.me/semantic-segmentation/#dilated_convolutions&quot;&gt;
&lt;div class=&quot;og-image&quot; style=&quot;background-image: url('https://scrap.kakaocdn.net/dn/cOsUsU/hyKWSW5JD6/CM2DYhWciDNFZkbCTZjj6k/img.png?width=2448&amp;amp;height=1780&amp;amp;face=0_0_2448_1780,https://scrap.kakaocdn.net/dn/BzvMg/hyKXSBboye/Gw0goJMcKmrBEeKm7tH0qk/img.png?width=2448&amp;amp;height=1780&amp;amp;face=0_0_2448_1780,https://scrap.kakaocdn.net/dn/erhIww/hyKWNO1CdY/wp8kWr4t4Xq9P7sMI6UA61/img.png?width=1194&amp;amp;height=693&amp;amp;face=0_0_1194_693');&quot;&gt;&amp;nbsp;&lt;/div&gt;
&lt;div class=&quot;og-text&quot;&gt;
&lt;p class=&quot;og-title&quot; data-ke-size=&quot;size16&quot;&gt;An overview of semantic image segmentation.&lt;/p&gt;
&lt;p class=&quot;og-desc&quot; data-ke-size=&quot;size16&quot;&gt;In this post, I'll discuss how to use convolutional neural networks for the task of semantic image segmentation. Image segmentation is a computer vision task in which we label specific regions of an image according to what's being shown. &quot;What's in this im&lt;/p&gt;
&lt;p class=&quot;og-host&quot; data-ke-size=&quot;size16&quot;&gt;www.jeremyjordan.me&lt;/p&gt;
&lt;/div&gt;
&lt;/a&gt;&lt;/figure&gt;
&lt;p data-ke-size=&quot;size16&quot;&gt;&amp;nbsp;&lt;/p&gt;</description>
      <category>인공지능/Segmentation</category>
      <author>통경</author>
      <guid isPermaLink="true">https://stat-cbc.tistory.com/29</guid>
      <comments>https://stat-cbc.tistory.com/29#entry29comment</comments>
      <pubDate>Wed, 21 Jul 2021 00:29:06 +0900</pubDate>
    </item>
    <item>
      <title>[풀잎스쿨 14기] semantic-segmenation-논문으로-입문하기 (U-Net_Elastic Deformations)</title>
      <link>https://stat-cbc.tistory.com/28</link>
      <description>&lt;p&gt;&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;b data-stringify-type=&quot;bold&quot;&gt;'본 포스팅은 모두의연구소(&lt;/b&gt;&lt;a href=&quot;http://home.modulabs.co.kr/&quot; data-stringify-link=&quot;http://home.modulabs.co.kr&quot; data-sk=&quot;tooltip_parent&quot;&gt;home.modulabs.co.kr&lt;/a&gt;&lt;span&gt;)&lt;/span&gt;&lt;b data-stringify-type=&quot;bold&quot;&gt;&lt;span&gt;&amp;nbsp;&lt;/span&gt;풀잎스쿨에서 진행된 '&lt;span style=&quot;color: #333333;&quot;&gt;semantic-segmenation&lt;/span&gt;' 과정 내용을 공유 및 정리한 자료입니다&lt;/b&gt;&lt;span&gt;.'&lt;/span&gt;&lt;/p&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;&lt;span&gt;1. introduction &lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;Sementic Segmentation 분야에서 가장 유명하다 할 수 있는 논문인 Unet paper(&lt;span&gt;&lt;/span&gt;&lt;a href=&quot;https://arxiv.org/abs/1505.04597&quot;&gt;https://arxiv.org/abs/1505.04597&lt;/a&gt;) 의 내용 중. 3.1 에서 Data Augmentation 에 관련하여 언급된 부분이 있었습니다.&amp;nbsp;&lt;/p&gt;
&lt;div data-ke-type=&quot;moreLess&quot; data-text-more=&quot;더보기&quot; data-text-less=&quot;닫기&quot;&gt;&lt;a class=&quot;btn-toggle-moreless&quot;&gt;더보기&lt;/a&gt;
&lt;div class=&quot;moreless-content&quot;&gt;
&lt;p&gt;We generate smooth deformations using random displacement vectors on a coarse 3 by 3 grid . The displacements are sampled from a Gaussian distribution with 10 pixels standard deviation.&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;해당 부분인데, U-net이란 논문 자체가 Biomedical Image Segmentation 에 중점을 두고 나온 논문이기 때문에 이 내용을 이해하기 위해 추가적인 의학 논문에서 해당 Elastic Deformations 기술을 사용한 논문이면서 자세하게 설명하고 있는 논문을 찾아 이를 기반으로 여러 자료를 추가하며 정리하였습니다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&amp;nbsp;&lt;/p&gt;
&lt;p&gt;Biomedical 분야에서 다양한 Deformations 방법 중 Elastic Deformations 을 사용한 이유는 다음과 같다.&amp;nbsp;&amp;nbsp;&lt;br /&gt;1) the small amount of available data&lt;/p&gt;
&lt;p&gt;2) class imbalance&amp;nbsp;&lt;/p&gt;
&lt;p&gt;이와 같은 문제를 해결하기 위해 elastic transformation을 도입하나, elastic transformation 이외에도 다른 augmentation도 충분히 사용 가능하다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&amp;gt;&amp;nbsp; Elastic Deformation 사용을 추천하는 경우 : 연속체에서 어떤 힘이나 시간 흐름으로 인해 변화가 발생하는 경우. 이 힘이 제거된 후 변형이 원래처럼 돌아오게 되면 이 변형을 탄성이라고 합니다. 이렇게 탄성이 있는 경우는 같은 물체라 해도 촬영 방법이나 각도 등에 의해서 다른 결과를 가져올 수 있으므로 이럴 때 사용하면 좋다고 하나 이 외의 경우에도 사용 (사람의 글씨체 차이에 따른 MNIST 적용 등에도 사용 했었음) 가능합니다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&amp;nbsp;&lt;/p&gt;
&lt;h3 data-ke-size=&quot;size23&quot;&gt;2. Method&lt;/h3&gt;
&lt;p&gt;1) 수평 및 수직 방향(x 및 y)에 대해 각각 임의의 stress (강도)가 generate 됨. 이 x,y 방향에 대한 강도를 다음과 같이 정의합니다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;$$\Delta_{x}, \Delta_{y}$$&lt;/p&gt;
&lt;p&gt;&amp;nbsp;&lt;/p&gt;
&lt;p&gt;각 픽셀 및 방향에 대해 다음과 같은 범위 안에서 생성됩니다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;$$\alpha \times [-0.5, 0.5] = [-0.5 \alpha, 0.5 \alpha]$$&lt;/p&gt;
&lt;p&gt;이 범위 안에서 uniformly 하게 선택됩니다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&amp;nbsp;&lt;/p&gt;
&lt;p&gt;아래의 가우스(가우시안) 필터(&lt;span&gt;&lt;/span&gt;&lt;a href=&quot;https://en.wikipedia.org/wiki/Gaussian_filter&quot;&gt;https://en.wikipedia.org/wiki/Gaussian_filter&lt;/a&gt;) 를 참조하시면 자세한 내용이 있습니다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;$$G(\sigma)= g(x,y)= \frac{1}{2 \pi \sigma^{2}} \cdot e^{-\frac{x^{2} + y^{2}}{2\sigma^{2}}}$$&lt;/p&gt;
&lt;p&gt;&amp;gt; x는 수평축 원점으로부터의 거리, y 는 수직축 원점으로부터의 거리, &amp;sigma; 는 가우스 분포의 표준 편차&lt;/p&gt;
&lt;p&gt;&amp;nbsp;&lt;/p&gt;
&lt;p&gt;2) 가까운 픽셀이 유사한 변위(displacement)를 갖도록하기 위해 결과 수평 및 수직 이미지에 Gaussian 필터를 별도로 적용합니다.&lt;/p&gt;
&lt;p&gt;$$\Delta_{x} = G(\sigma) * (\alpha \times Rand(n,m))$$&lt;/p&gt;
&lt;p&gt;$$\Delta_{y} = G(\sigma) * (\alpha \times Rand(n,m))$$&lt;/p&gt;
&lt;p&gt;&amp;nbsp;&lt;/p&gt;
&lt;p&gt;이러한 변환에는 두 가지 매개 변수가 있습니다.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;무작위 초기 변위에서의 최댓값 (alpha)&lt;/li&gt;
&lt;li&gt;가우시안 필터의 표준 편차에 의해 주어진 평활화 작업의 강도(sigma)&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;결과 패치 모양에 따라 이 값을 alpha = 300, sigma = 20 으로 설정. 그 후, stress 필드가 이미지, breast segmentation 마스크 및 Mass annotation 에 적용됩니다. 이것은 각 픽셀을 새로운 위치로 이동하고 정수 좌표에서 강도를 얻기 위해 spline interpolation 을 사용하여 수행됩니다.&lt;/p&gt;
&lt;p&gt;$$I_{trans}(j + \Delta_{x}(j,k), k + \Delta_{y}(j,k)) = I(j,k)$$&lt;/p&gt;
&lt;p&gt;위의 방정식에서 I와 Itrans는 각각 원본 이미지와 변환된 이미지입니다. 그리고 n * m 은 image dimensions입니다.&lt;/p&gt;
&lt;p&gt;이러한 변형의 영향에 대한 예 그림은 아래와 같으며, 이는 유방암에 관련한 참조 연구 논문에서 발췌하였습니다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-filename=&quot;Untitled.png&quot; data-origin-width=&quot;605&quot; data-origin-height=&quot;434&quot; data-ke-mobilestyle=&quot;widthContent&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/bO8el4/btq0AHTZ8LM/lqIRCnWnYdVXrqXGGhK6m1/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/bO8el4/btq0AHTZ8LM/lqIRCnWnYdVXrqXGGhK6m1/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/bO8el4/btq0AHTZ8LM/lqIRCnWnYdVXrqXGGhK6m1/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FbO8el4%2Fbtq0AHTZ8LM%2FlqIRCnWnYdVXrqXGGhK6m1%2Fimg.png&quot; data-filename=&quot;Untitled.png&quot; data-origin-width=&quot;605&quot; data-origin-height=&quot;434&quot; data-ke-mobilestyle=&quot;widthContent&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;&amp;nbsp;&lt;/p&gt;
&lt;p&gt;U-Net 예제에서 나온 Cell에 관련된 적용 이미지는 다음과 같습니다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;512 * 512 왼쪽 원본 이미지 변형으로 U-Net 에서 사용한 Elastic Deformations 변형을 적용한 것이 오른쪽입니다.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;figure class=&quot;imageblock alignCenter&quot; data-filename=&quot;Untitled2.png&quot; data-origin-width=&quot;640&quot; data-origin-height=&quot;480&quot; data-ke-mobilestyle=&quot;widthContent&quot;&gt;&lt;span data-url=&quot;https://blog.kakaocdn.net/dn/nGTnu/btq0xJ6Q5b7/mfSSPt0hojYronkBUIzH11/img.png&quot; data-phocus=&quot;https://blog.kakaocdn.net/dn/nGTnu/btq0xJ6Q5b7/mfSSPt0hojYronkBUIzH11/img.png&quot;&gt;&lt;img src=&quot;https://blog.kakaocdn.net/dn/nGTnu/btq0xJ6Q5b7/mfSSPt0hojYronkBUIzH11/img.png&quot; srcset=&quot;https://img1.daumcdn.net/thumb/R1280x0/?scode=mtistory2&amp;fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FnGTnu%2Fbtq0xJ6Q5b7%2FmfSSPt0hojYronkBUIzH11%2Fimg.png&quot; data-filename=&quot;Untitled2.png&quot; data-origin-width=&quot;640&quot; data-origin-height=&quot;480&quot; data-ke-mobilestyle=&quot;widthContent&quot; onerror=&quot;this.onerror=null; this.src='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png'; this.srcset='//t1.daumcdn.net/tistory_admin/static/images/no-image-v1.png';&quot;/&gt;&lt;/span&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;Elastic Deformations 에 대해 알아보았습니다!&amp;nbsp; 이렇게 논문을 심층적으로 분석하고 함께 알아가는 풀잎스쿨 15기가 모집중이니 어서 풀잎스쿨 열차에 탑승해보세요!&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://home.modulabs.co.kr/apply-flip15th/&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot;&gt;home.modulabs.co.kr/apply-flip15th/&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&amp;nbsp;&lt;/p&gt;
&lt;p&gt;&amp;nbsp;&lt;/p&gt;
&lt;p&gt;References.&amp;nbsp;&lt;/p&gt;
&lt;p&gt;Elastic Deformations for Data Augmentation in Breast Cancer Mass Detection, Eduardo Castro and Jaime S. Cardoso and J. C. Pereira, 2018 IEEE EMBS International Conference on Biomedical &amp;amp; Health Informatics (BHI), 230-234p U-Net: Convolutional Networks for Biomedical Image Segmentation, Olaf Ronneberger and Philipp Fischer and Thomas Brox, 2015&lt;/p&gt;</description>
      <category>인공지능/Computer Vision</category>
      <category>DataAugmentation</category>
      <category>sementicsegmentation</category>
      <category>UNET</category>
      <category>Vision</category>
      <category>데이터</category>
      <category>데이터분석</category>
      <category>딥러닝</category>
      <category>비전</category>
      <category>유넷</category>
      <category>풀잎스쿨14기</category>
      <author>통경</author>
      <guid isPermaLink="true">https://stat-cbc.tistory.com/28</guid>
      <comments>https://stat-cbc.tistory.com/28#entry28comment</comments>
      <pubDate>Sat, 20 Mar 2021 14:51:14 +0900</pubDate>
    </item>
  </channel>
</rss>